Asked what the distinction actually is, and the SQL half needed correcting first: under exclusive ownership SQL runs against your own database and nothing else, whatever transport a query might travel over. Both options are the mesh's own channel and both ride the broker, so the transport is not the distinction. The distinction is where the answer lives when you need it. A request asks at the moment and waits — always current, costs a round trip, cannot answer when the other side is down. A subscription keeps a local copy — instant, works offline, as current as the last event received, and you must handle what you missed. What decides is not taste. ADR 0036 makes disconnection an ordinary situation rather than an exception, so anything that must keep working while disconnected CANNOT use a request: there is nobody to ask. And the converse — anything where a stale answer is worse than no answer cannot use a subscription. A display can lag; a decision about whether a grant is still valid cannot. So an apparently open question turns out to be derived from a decision already taken. What stays open is narrower: what a consumer does about the events it missed while disconnected — replay from a point, ask once for a full picture and resume, or rebuild. The question every projection has. Also recorded: separate databases are required in the new design, and the shared registry is a leftover rather than a pattern.
454 lines
24 KiB
Markdown
454 lines
24 KiB
Markdown
# A provider, all the way through
|
|
|
|
One module, worked out completely, because it is the case that breaks the tidy version. It
|
|
looks like *a container that runs a database* and it is at least nine things.
|
|
|
|
Postgres is the example. **The shape is not specific to it** — see the end.
|
|
|
|
## What it carries
|
|
|
|
**1 — A supervised container.** An image, a version, a data volume, and configuration. Easy,
|
|
and the only part the phrase "a docker service" describes.
|
|
|
|
**2 — Persistent state, and where it lives matters.** The volume is the database. Moving this
|
|
module between nodes is not rescheduling; it is a migration. Almost nothing else in the
|
|
catalogue has this property, and nothing in `provides` / `requires` expresses it.
|
|
|
|
**3 — Configuration that is partly the machine's.** Tuning follows the hardware — memory,
|
|
storage. A declaration generated centrally cannot know those, and
|
|
[research 012](../012-the-minimum-viable-node/00-overview.md) says the machine's own values win
|
|
on conflict. So some of this module's configuration is *derived from the node it lands on*.
|
|
|
|
**4 — A tool surface.** It exposes capabilities agents can call — query, list, provision. That
|
|
is not a resource on a machine and not an artifact; it is a contract the mesh publishes on the
|
|
module's behalf.
|
|
|
|
**5 — A provisioner.** The part that matters, and the one below.
|
|
|
|
**6 — Its own bookkeeping.** The provisioner must remember what it granted to whom, or it
|
|
cannot revoke, rotate or clean up. So a module that provides state to others *also* holds state
|
|
about its providing — and that state is not the database's data.
|
|
|
|
**7 — An exposure decision, per node it runs on.** Reachable from the machine only, from the
|
|
local network, or publicly. That is a property of *this assignment*, not of the module: the
|
|
same module on two nodes may answer differently.
|
|
|
|
**8 — Credentials it generates.** Per consumer, and they have to reach the consumer. Which
|
|
means a provisioning edge carries a payload, and the payload is a secret.
|
|
|
|
**9 — Health that is not "the container is up".** A container running and a database accepting
|
|
connections are different facts, and the second is the one anything cares about. This is the
|
|
host's read-back rule, at a distance.
|
|
|
|
## The provisioner is a second kind of edge
|
|
|
|
The tidy version of this effort said: *a module provides names, a module requires names, that is
|
|
the only edge.* A small game wanting to store data shows it is not.
|
|
|
|
```
|
|
my-cool-game requires postgres # I speak its protocol, it must exist
|
|
my-cool-game requires a database FROM postgres, called my-cool-game
|
|
```
|
|
|
|
The first is **presence**: the thing exists and is reachable. Nothing is created; nothing flows
|
|
back. `vscode requires terminal` is this, and so is `requires container-runtime`.
|
|
|
|
The second is **instantiation**: the provider is asked to make something *for this consumer*,
|
|
and hands back what the consumer needs to use it. A database, a user, a password, an address.
|
|
|
|
They differ in every way that matters:
|
|
|
|
| | presence | instantiation |
|
|
|---|---|---|
|
|
| creates something | no | yes, one per consumer |
|
|
| carries a payload back | no | credentials, an address |
|
|
| can be revoked | — | yes, and must be when the consumer goes |
|
|
| provider holds state about it | no | yes — who was granted what |
|
|
| satisfied by | anything providing the name | that provider, specifically |
|
|
|
|
**This is the mesh's actual power**, in the operator's words: a small game declares it wants a
|
|
database and the mesh makes one. Nobody creates a user by hand, nobody pastes a connection
|
|
string. That is worth being precise about rather than folding into a single relation because
|
|
one relation is prettier.
|
|
|
|
## What that costs the design
|
|
|
|
**The proposal's "one kind of edge" is wrong**, and the current system already knew: it has
|
|
`dependencies` for presence and `requires: provision:` for instantiation, with the resolver
|
|
deriving a presence edge from every instantiation edge. [`analysis.md`](analysis.md) recorded
|
|
that derivation as a convenience. It is not — it is the correct relationship between two
|
|
genuinely different relations.
|
|
|
|
So: **two kinds of edge, one graph.** Instantiation implies presence. Presence does not imply
|
|
instantiation.
|
|
|
|
## Which provider, when two nodes run one
|
|
|
|
Asked directly, because two nodes can each run a relational store and a consumer has to be
|
|
served by one of them. Neither obvious answer is right.
|
|
|
|
**Not "the consumer names the node."** That is placement in the consumer's manifest — a small
|
|
game edited because a database moved, which is the fault
|
|
[`proposal.md`](proposal.md) separates constraints from placement to avoid.
|
|
|
|
**Not "the consumer does not care" either.** For presence it genuinely does not: a terminal is a
|
|
terminal. For instantiation it cares permanently, because the data lands in exactly one store
|
|
and the wrong choice is discovered long afterwards.
|
|
|
|
**What the consumer does know is the scope of its own need.** Not which node — how many of the
|
|
thing it wants, relative to itself:
|
|
|
|
| Scope | Means | Example |
|
|
|---|---|---|
|
|
| **shared** | one instance serves every instance of this consumer | the mesh's own registry: every node reads the same rows |
|
|
| **per instance** | each instance of this consumer gets its own | a local cache, a per-node queue |
|
|
|
|
That is a property of the consumer, expressible without naming anything. And it is the thing
|
|
that actually decides: a shared need cannot be satisfied by a provider each node runs
|
|
separately, and a per-instance need should not be satisfied by a shared one.
|
|
|
|
**Then the mesh binds, and the binding is written down.** Not recomputed: a resolver that
|
|
re-derives which store serves a consumer will one day derive a different answer and relocate a
|
|
database, so the binding is made once and changed only deliberately.
|
|
|
|
**Where it is written down follows the scope.** A shared grant belongs to the *module* and every
|
|
assignment of it references the same one — which is the answer for one module installed on two
|
|
nodes wanting one database between them. A per-instance grant belongs to the *assignment*. Same
|
|
relation, two homes, and which home is not a detail: it is what makes two instances share
|
|
something or not.
|
|
|
|
The pieces that follow, none of them settled here:
|
|
|
|
- **When several providers satisfy the scope**, something chooses — most plausibly locality,
|
|
preferring a provider on the same node. That is a default, and it must be overridable, because
|
|
the reason to override it is exactly the reason nobody anticipated it.
|
|
- **A binding is a thing that can be wrong.** Once recorded it can be inspected, and a consumer
|
|
bound to a store on a node that no longer exists is a question somebody can be asked rather
|
|
than a failure at connect time.
|
|
- **Moving a binding moves data.** Whatever the mechanism, changing it is a migration and not a
|
|
configuration change, and a design that lets it look like the latter will lose something.
|
|
|
|
## Provisioning is early, not late
|
|
|
|
An assumption worth killing: that provisioning is something the control plane does for
|
|
consumers once a mesh is running.
|
|
|
|
**The mesh's own registry is a provisioned database.** So is its virtual host on the broker.
|
|
Neither exists until something creates them, and nothing in the mesh works until they do. The
|
|
order is:
|
|
|
|
```
|
|
1 the store runs from the bundle the host carries
|
|
2 a database is created in it a provisioning step
|
|
3 the mesh's own schema is applied a migration, against that database
|
|
4 the control plane starts and only now is there a mesh
|
|
5 everything else is provisioned the ordinary path
|
|
```
|
|
|
|
Steps 2 and 3 happen **before there is a mesh to do them**. So provisioning is not a
|
|
control-plane service that consumers use; it is part of the bootstrap, and part of what the
|
|
carried bundle has to be able to express.
|
|
|
|
**Which strains what a declaration is.** [ADR 0043](../../02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md)
|
|
has the host applying *declared state on this machine*. A database inside a running store is not
|
|
a file or a unit — and at bootstrap it is, at least, local: the store is on the same machine as
|
|
the host applying the bundle.
|
|
|
|
Later it is not. A consumer on one node provisioned from a store on another is the ordinary
|
|
case, and reaching it is not the host's job. So the same operation is local at bootstrap and
|
|
remote afterwards, which is either two mechanisms or one mechanism with a boundary crossing in
|
|
it. Undecided, and it is the sharpest unresolved thing in this file.
|
|
|
|
## Several modules, one database — and the case for refusing
|
|
|
|
Asked, then reconsidered by the operator: *maybe we should not allow it.*
|
|
|
|
The permissive version was a per-consumer **schema** inside a shared database — its own
|
|
namespace, its own migrations, revocable by dropping the schema, with a cross-context join
|
|
possible but deliberate.
|
|
|
|
**The stricter version is better, and it goes further than schemas.**
|
|
|
|
> **A module is only ever granted a resource it exclusively owns.**
|
|
|
|
No shared writes. And **no read-only role on another module's database either** — reading
|
|
another context's tables couples you to its layout exactly as firmly as writing them does, and
|
|
the coupling is harder to see because nothing breaks until the owner changes a column.
|
|
|
|
That is what [`how-we-build`](../../00-META/how-we-build.md) §4 already says: *contexts integrate
|
|
through the record, never through a shared schema.* The permissive version kept the letter of it
|
|
and left the temptation in place, and the path of least resistance wins eventually. A boundary
|
|
that is merely inconvenient to cross is a boundary that gets crossed.
|
|
|
|
### What it costs
|
|
|
|
**Cross-module reporting.** Anything wanting to know what several modules hold can no longer
|
|
join across them. It consumes their events, or calls their interface, and neither is as
|
|
immediate as a query.
|
|
|
|
That cost is the point rather than a regrettable side effect — it is §4's whole argument, and
|
|
the mesh already has both mechanisms: an event stream every context publishes to, and a tool
|
|
surface every module exposes. What gets harder is the thing that was making work belonging to
|
|
one context keep having to be implemented in another.
|
|
|
|
**One more connection per consumer.** A dozen modules means a dozen databases rather than a
|
|
dozen schemas in one. For a relational store this is unremarkable; it is worth stating only so
|
|
nobody discovers it as a surprise.
|
|
|
|
### What it deletes
|
|
|
|
The effort has been looking for what the design *removes* rather than adds, and this is the
|
|
first clear instance:
|
|
|
|
- **Grant kinds.** There is one — an exclusive resource. No schema grants, no read roles, no
|
|
scoping rules for who may see what inside a shared thing.
|
|
- **The question of who owns which table**, and with it the guessing at revocation time.
|
|
- **Cross-module migration ordering.** Two modules migrating one database need their migrations
|
|
ordered against each other. Exclusive ownership means a module's migrations are ordered only
|
|
against itself.
|
|
- **A whole class of permission modelling** that a shared store would otherwise need.
|
|
|
|
### The hardest case for it: a dashboard reading everything
|
|
|
|
Raised immediately, and it is the right test — a board showing nodes, modules, pipelines, agents
|
|
and tasks wants to read a dozen stores. Under this rule it cannot read any of them.
|
|
|
|
**It survives, and not by luck: the board is a surface, and surfaces already may not do this.**
|
|
The skeleton puts `api/` in the control plane as *the one interface every surface speaks to*, and
|
|
tier 3 as *thin; no logic lives here*. A board reading stores directly is a surface reaching past
|
|
the context that owns the data — which the tier rule forbids for reasons that have nothing to do
|
|
with databases.
|
|
|
|
So the dashboard is not a counter-example. It is an instance the rule catches, and the two rules
|
|
turn out to be one rule seen from two sides: **exclusive ownership is the tier rule, expressed in
|
|
terms of storage.**
|
|
|
|
**The general shape, for anything that needs to see across many things:** consume the record and
|
|
own your own view. A reporting context builds a projection from events and reads its own store.
|
|
It does not read anybody else's.
|
|
|
|
**And the cost is not small, so it should be said plainly.** A projection is more work than a
|
|
join, and it lags. Today a board queries the mesh's own database directly — that is ordinary,
|
|
it works, and this rule makes it a migration rather than a preference. The reason to pay it is
|
|
[`how-we-build`](../../00-META/how-we-build.md) §4's already-measured cost, not elegance.
|
|
|
|
### Checked against the real consumers
|
|
|
|
The rule's survival turned on whether every reader of the mesh's own registry could be served
|
|
some other way. **Eighteen consumers open a direct connection to it.** Four groups, and only one
|
|
of them is work.
|
|
|
|
| Group | What happens under the rule |
|
|
|---|---|
|
|
| **The owner and its machinery** — the mesh module, the SDK, the environment and configuration synchronisers, secrets | Nothing. It owns the database. |
|
|
| **Node appliers** — the overlay, the shell daemon, the resolver | **Already resolved.** [ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md) stops the host querying the mesh database, decided for tier reasons with nothing to do with this. |
|
|
| **Foreign tenants** — the work engine (10 tables), the knowledge base (2), pipeline logs (1) | They need **their own database**. They are not reading the registry; they are storing their own data in it. |
|
|
| **Genuine cross-context reads** — the work engine reads `nodes`; two others read a handful | The only ones needing an interface or events. |
|
|
|
|
**Thirteen foreign tables live in the mesh's registry database**, belonging to three separate
|
|
contexts. That is [`how-we-build`](../../00-META/how-we-build.md) §4's shared schema, counted.
|
|
|
|
**So the question was the wrong shape.** The bulk of the problem is not readers needing a new
|
|
route to data — it is **tenants needing to move out**. Tasks, agents and teams have nothing to do
|
|
with nodes and modules; they are co-located by history. Give that context its own database and
|
|
its dependency on the registry shrinks to a single table.
|
|
|
|
What remains is a handful of genuine cross-context reads, small enough to enumerate rather than
|
|
estimate. **The rule holds.**
|
|
|
|
### Request or subscription, and what decides
|
|
|
|
The remaining cross-context reads need one or the other. **Neither is SQL** — under exclusive
|
|
ownership a module runs SQL against its own database and nothing else, whatever transport a
|
|
query might travel over. Both options are the mesh's own channel, and both ride the broker
|
|
([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md)), so the transport is
|
|
not the distinction.
|
|
|
|
**The distinction is where the answer lives when you need it.**
|
|
|
|
| | request | subscription |
|
|
|---|---|---|
|
|
| you ask | at the moment you need to know | never — you are told |
|
|
| the answer lives | on the other side | in your own store |
|
|
| freshness | always current | as current as the last event you received |
|
|
| when the other side is down | you cannot answer | you answer from your copy |
|
|
| what you must handle | a round trip that can fail | events you missed while you were down |
|
|
|
|
**What decides is not taste.** [ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)
|
|
makes disconnection an ordinary situation rather than an exception. So:
|
|
|
|
> **Anything that must keep working while disconnected cannot use a request** — there is nobody
|
|
> to ask. It needs a local copy, which means a subscription.
|
|
|
|
And the converse: anything that must be *correct at the instant of asking*, where a stale answer
|
|
is worse than no answer, cannot use a subscription. A display can lag. A decision about whether
|
|
a grant is still valid cannot.
|
|
|
|
That turns an apparently open question into a derived one. What remains genuinely open is
|
|
narrower: **what a consumer does about the events it missed** while it was disconnected — replay
|
|
from a point, ask once for a full picture and resume, or rebuild from scratch. That is the same
|
|
question every projection has, and nothing in the record answers it yet.
|
|
|
|
## A migration belongs to the consumer and runs on the provider
|
|
|
|
A game defines migrations. They run against the database the store granted **it**. So the
|
|
migration is:
|
|
|
|
- **owned** by the consumer — it is that module's schema, versioned with that module;
|
|
- **hosted** by the provider — it runs inside something the consumer does not control;
|
|
- **ordered** after the provisioning edge — there is nothing to migrate until the grant exists;
|
|
- **scoped** to the grant — the consumer's migrations touch its database and no other.
|
|
|
|
Ownership crosses the edge, which nothing in *provides* and *requires* expresses. And it is the
|
|
*action* category from [`features.md`](features.md) made concrete: not an artifact, not node
|
|
state, and not something the host can apply, because the thing it changes is not the machine.
|
|
|
|
It also gives a consumer's own install an internal order — **provisioned, then migrated, then
|
|
started** — that depends on an edge rather than on the module's contents.
|
|
|
|
## What still has no answer
|
|
|
|
**How many instances of postgres should exist?** One per mesh is wrong — a node that must work
|
|
while disconnected cannot depend on a database elsewhere. One per node is wrong — the mesh's own
|
|
registry is one thing, not one per node. So the answer is per-module, and nothing in the schema
|
|
says it. This is [`cases.md`](cases.md) axis *how many instances*, and postgres is the case that
|
|
proves it cannot be a global rule.
|
|
|
|
**What happens to a grant when the consumer is removed?** The game is uninstalled. Its database
|
|
still exists, holding its data. Dropping it silently is data loss; keeping it forever is a leak.
|
|
[ADR 0043](../../02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md) says
|
|
the host removes what it applied and no longer declares — but this is not on the host, it is
|
|
inside another module's state, and the same reasoning does not obviously carry.
|
|
|
|
**Where does node-derived configuration come from?** (3) The control plane composes a
|
|
declaration, and cannot know this machine's memory. Either the host fills in a blank the
|
|
declaration leaves — which makes the host decide something, against
|
|
[ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md) — or the control
|
|
plane reads the node's inventory first and composes with it. The second is consistent and means
|
|
a declaration is composed *per node from what the node reported*, which is a stronger claim than
|
|
anything recorded so far.
|
|
|
|
|
|
## Providing is not a substrate thing
|
|
|
|
The four pinned services are the obvious providers, and they are not a category.
|
|
|
|
| Service | What a consumer asks it for |
|
|
|---|---|
|
|
| a relational store | a database, a user, credentials |
|
|
| another relational store, different vendor | a database — **and not the same one** |
|
|
| a message broker | a virtual host, a user, permissions |
|
|
| an object store | a bucket and keys |
|
|
| an image registry | a repository |
|
|
| an identity provider | a client, a realm, a secret |
|
|
| an analytics service | a site, and a tracking identity |
|
|
| a low-code data platform | a base, and a token |
|
|
| an application platform | a project, which is several of the above at once |
|
|
| a mail server | a mailbox, an alias, credentials |
|
|
|
|
**Any hosted service can be a factory.** Providing is a facet a module may have, not a kind of
|
|
module it is — which is the same conclusion the effort reached about services and applications,
|
|
arriving from the other direction.
|
|
|
|
That kills the last reason to keep *provider* as a category. A module runs something, or grants
|
|
something, or both, or neither.
|
|
|
|
### Two stores, and why `database` still is not a name
|
|
|
|
Two relational stores from different vendors both grant *a database*. They are the sharpest
|
|
possible test of the substitutability rule from [`proposal.md`](proposal.md), and they fail it
|
|
completely: different wire protocol, different dialect, different driver, different client
|
|
library compiled into the consumer.
|
|
|
|
A consumer declaring `requires: database` and being handed either would break against one of
|
|
them. So the name promises what no provider delivers — and now with two real providers in the
|
|
catalogue rather than a thought experiment.
|
|
|
|
*Database* remains a **tag**. It is how a person finds both. It is not how a consumer names what
|
|
it needs.
|
|
|
|
## The same shape, three more times
|
|
|
|
The message broker has all nine. So does the object store, and so does the image registry. They
|
|
differ in what a consumer asks for — a database, a virtual host, a bucket, a repository — and in
|
|
nothing structural.
|
|
|
|
**A substrate service is a service plus a factory.** That is the whole pattern, and there are
|
|
four of them. It generalises past the substrate too: anything that grants something per consumer
|
|
has this shape, and anything that does not is the simpler case.
|
|
|
|
But two things differ *between* them, and both matter more than the similarity.
|
|
|
|
### The broker cannot be managed over the broker
|
|
|
|
[ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md) makes the broker the
|
|
channel every node takes work from, and
|
|
[ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md) makes it the security
|
|
boundary — everything a node applies arrives through it.
|
|
|
|
So the module providing the broker is also **the way modules are managed**. A declaration cannot
|
|
be delivered to it over itself, and reconfiguring it is done through the thing being
|
|
reconfigured. Nothing else in the catalogue has that property; the store is consumed by the
|
|
control plane but is not how the control plane *reaches* anything.
|
|
|
|
This is exactly what the carried bundle exists for
|
|
([ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md)): the broker is raised from
|
|
what the host carries, before there is a channel, because there is no other way to raise it.
|
|
Recorded here because it is a constraint on *one module*, not a general rule, and a schema with
|
|
no way to say so hides it.
|
|
|
|
### Two modules of identical shape want different instance counts
|
|
|
|
The broker is one per mesh — a single point of failure and a single point of trust, by decision
|
|
rather than by accident. The store cannot be: a node that must keep working while disconnected
|
|
([ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)) cannot depend on a database
|
|
somewhere else.
|
|
|
|
Same nine properties, opposite answers. Which settles something the cases file left open: **how
|
|
many instances is not derivable from what a module is.** It is a decision per module, it has to
|
|
be declared, and nothing in `provides`, `requires` or `excludes` says it.
|
|
|
|
### And revocation differs in consequence
|
|
|
|
Revoking a database leaves data behind until something drops it — a leak, and recoverable.
|
|
Revoking a virtual host drops whatever had not been delivered — not recoverable, and silent.
|
|
|
|
The relation is the same and the blast radius is not, which is an argument for the provider
|
|
deciding what revocation means rather than the mesh applying one rule to all of them.
|
|
|
|
|
|
## The assignment is a third thing
|
|
|
|
Recorded because the operator tried the alternative and abandoned it: **modules were once
|
|
node-agnostic**, and it did not survive contact.
|
|
|
|
The worked example says why. Several of a provider's nine properties are not properties of the
|
|
module at all:
|
|
|
|
- **where its state lives** — a volume on a particular machine;
|
|
- **how it is reached** — the same module on two nodes may answer locally, on the network, or
|
|
publicly, and that is a per-assignment decision;
|
|
- **configuration derived from the hardware** — tuning follows the memory and storage of the
|
|
machine it landed on;
|
|
- **whether this instance is the one** a given consumer is provisioned from.
|
|
|
|
None of those belong in the catalogue, because they differ per node. None belong in the node,
|
|
because they are about this module. **They belong to the pairing**, and a design with only
|
|
modules and nodes has nowhere to put them — which is what "node-agnostic" ran out of.
|
|
|
|
So there are three entities, not two:
|
|
|
|
> **a module** · **a node** · **an assignment**, which is a module on a node and carries its own
|
|
> configuration
|
|
|
|
The current system already has this, arrived at the same way: environment values are stored per
|
|
module *and per node*, so a module's settings differ between the machines running it.
|
|
|
|
**This does not put placement back in the manifest.** A module still says what must be true of a
|
|
node and never which node ([`proposal.md`](proposal.md)). What changes is that the *result* of
|
|
placing it is a thing with its own state, rather than a fact recorded on one of the two ends.
|
|
|
|
And it makes the composed declaration question from above answerable: a declaration is built
|
|
from the module, the node's inventory, and the assignment between them. Three inputs, which is
|
|
why two were never enough.
|