Files
hq/01-RESEARCH/011-the-module-graph/worked-provider.md
T
jschoubben e1febe8e0f Renumber the records 1 to 23
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18,
19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only
the archaeology of what used to be there.

Renumbered contiguously. Renames run in ascending order, so every target number
is already free and no two files ever collide.

The reference rewrite is one simultaneous pass rather than a sequence of
replacements. Numbers moved into slots other numbers were vacating -- the node
host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time
would have cascaded and silently pointed things at the wrong record.

Seven plain-text references survived the merges as prose rather than links,
naming records that no longer existed: the enrolment token, the link boundary,
what a declaration is, reachability, the repository structure. Each mapped to
the consolidated record that now holds it.

Verified rather than assumed: every [ADR NNNN](path) link now has matching text
and target, checked across the whole repository, and the checker passes.

Frontmatter `consolidates:` lists dropped -- they named records that are gone,
and each consolidated record already says in prose what it absorbed.
2026-08-28 23:28:34 +02:00

26 KiB

A provider, all the way through

One module, worked out completely, because it is the case that breaks the tidy version. It looks like a container that runs a database and it is at least nine things.

Postgres is the example. The shape is not specific to it — see the end.

What it carries

1 — A supervised container. An image, a version, a data volume, and configuration. Easy, and the only part the phrase "a docker service" describes.

2 — Persistent state, and where it lives matters. The volume is the database. Moving this module between nodes is not rescheduling; it is a migration. Almost nothing else in the catalogue has this property, and nothing in provides / requires expresses it.

3 — Configuration that is partly the machine's. Tuning follows the hardware — memory, storage. A declaration generated centrally cannot know those, and research 012 says the machine's own values win on conflict. So some of this module's configuration is derived from the node it lands on.

4 — A tool surface. It exposes capabilities agents can call — query, list, provision. That is not a resource on a machine and not an artifact; it is a contract the mesh publishes on the module's behalf.

5 — A provisioner. The part that matters, and the one below.

6 — Its own bookkeeping. The provisioner must remember what it granted to whom, or it cannot revoke, rotate or clean up. So a module that provides state to others also holds state about its providing — and that state is not the database's data.

7 — An exposure decision, per node it runs on. Reachable from the machine only, from the local network, or publicly. That is a property of this assignment, not of the module: the same module on two nodes may answer differently.

8 — Credentials it generates. Per consumer, and they have to reach the consumer. Which means a provisioning edge carries a payload, and the payload is a secret.

9 — Health that is not "the container is up". A container running and a database accepting connections are different facts, and the second is the one anything cares about. This is the host's read-back rule, at a distance.

The provisioner is a second kind of edge

The tidy version of this effort said: a module provides names, a module requires names, that is the only edge. A small game wanting to store data shows it is not.

my-cool-game   requires   postgres            # I speak its protocol, it must exist
my-cool-game   requires   a database FROM postgres, called my-cool-game

The first is presence: the thing exists and is reachable. Nothing is created; nothing flows back. vscode requires terminal is this, and so is requires container-runtime.

The second is instantiation: the provider is asked to make something for this consumer, and hands back what the consumer needs to use it. A database, a user, a password, an address.

They differ in every way that matters:

presence instantiation
creates something no yes, one per consumer
carries a payload back no credentials, an address
can be revoked — yes, and must be when the consumer goes
provider holds state about it no yes — who was granted what
satisfied by anything providing the name that provider, specifically

This is the mesh's actual power, in the operator's words: a small game declares it wants a database and the mesh makes one. Nobody creates a user by hand, nobody pastes a connection string. That is worth being precise about rather than folding into a single relation because one relation is prettier.

What that costs the design

The proposal's "one kind of edge" is wrong, and the current system already knew: it has dependencies for presence and requires: provision: for instantiation, with the resolver deriving a presence edge from every instantiation edge. analysis.md recorded that derivation as a convenience. It is not — it is the correct relationship between two genuinely different relations.

So: two kinds of edge, one graph. Instantiation implies presence. Presence does not imply instantiation.

Which provider, when two nodes run one

Asked directly, because two nodes can each run a relational store and a consumer has to be served by one of them. Neither obvious answer is right.

Not "the consumer names the node." That is placement in the consumer's manifest — a small game edited because a database moved, which is the fault proposal.md separates constraints from placement to avoid.

Not "the consumer does not care" either. For presence it genuinely does not: a terminal is a terminal. For instantiation it cares permanently, because the data lands in exactly one store and the wrong choice is discovered long afterwards.

What the consumer does know is the scope of its own need. Not which node — how many of the thing it wants, relative to itself:

Scope Means Example
shared one instance serves every instance of this consumer the mesh's own registry: every node reads the same rows
per instance each instance of this consumer gets its own a local cache, a per-node queue

That is a property of the consumer, expressible without naming anything. And it is the thing that actually decides: a shared need cannot be satisfied by a provider each node runs separately, and a per-instance need should not be satisfied by a shared one.

Then the mesh binds, and the binding is written down. Not recomputed: a resolver that re-derives which store serves a consumer will one day derive a different answer and relocate a database, so the binding is made once and changed only deliberately.

Where it is written down follows the scope. A shared grant belongs to the module and every assignment of it references the same one — which is the answer for one module installed on two nodes wanting one database between them. A per-instance grant belongs to the assignment. Same relation, two homes, and which home is not a detail: it is what makes two instances share something or not.

The pieces that follow, none of them settled here:

  • When several providers satisfy the scope, something chooses — most plausibly locality, preferring a provider on the same node. That is a default, and it must be overridable, because the reason to override it is exactly the reason nobody anticipated it.
  • A binding is a thing that can be wrong. Once recorded it can be inspected, and a consumer bound to a store on a node that no longer exists is a question somebody can be asked rather than a failure at connect time.
  • Moving a binding moves data. Whatever the mechanism, changing it is a migration and not a configuration change, and a design that lets it look like the latter will lose something.

Provisioning is early, not late

An assumption worth killing: that provisioning is something the control plane does for consumers once a mesh is running.

The mesh's own registry is a provisioned database. So is its virtual host on the broker. Neither exists until something creates them, and nothing in the mesh works until they do. The order is:

1  the store runs                       from the bundle the host carries
2  a database is created in it          a provisioning step
3  the mesh's own schema is applied     a migration, against that database
4  the control plane starts             and only now is there a mesh
5  everything else is provisioned       the ordinary path

Steps 2 and 3 happen before there is a mesh to do them. So provisioning is not a control-plane service that consumers use; it is part of the bootstrap, and part of what the carried bundle has to be able to express.

Which strains what a declaration is. ADR 0016 has the host applying declared state on this machine. A database inside a running store is not a file or a unit — and at bootstrap it is, at least, local: the store is on the same machine as the host applying the bundle.

Later it is not. A consumer on one node provisioned from a store on another is the ordinary case, and reaching it is not the host's job.

Resolved as two mechanisms, which is the answer rather than a compromise (ADR 0016). The host runs bootstrap actions locally from the bundle; the control plane provisions across the mesh afterwards. Different actors, different scopes, different trust paths — so there is no single operation with a tier boundary running through it.

Several modules, one database — and the case for refusing

Asked, then reconsidered by the operator: maybe we should not allow it.

The permissive version was a per-consumer schema inside a shared database — its own namespace, its own migrations, revocable by dropping the schema, with a cross-context join possible but deliberate.

The stricter version is better, and it goes further than schemas.

A module is only ever granted a resource it exclusively owns.

No shared writes. And no read-only role on another module's database either — reading another context's tables couples you to its layout exactly as firmly as writing them does, and the coupling is harder to see because nothing breaks until the owner changes a column.

That is what how-we-build §4 already says: contexts integrate through the record, never through a shared schema. The permissive version kept the letter of it and left the temptation in place, and the path of least resistance wins eventually. A boundary that is merely inconvenient to cross is a boundary that gets crossed.

What it costs

Cross-module reporting. Anything wanting to know what several modules hold can no longer join across them. It consumes their events, or calls their interface, and neither is as immediate as a query.

That cost is the point rather than a regrettable side effect — it is §4's whole argument, and the mesh already has both mechanisms: an event stream every context publishes to, and a tool surface every module exposes. What gets harder is the thing that was making work belonging to one context keep having to be implemented in another.

One more connection per consumer. A dozen modules means a dozen databases rather than a dozen schemas in one. For a relational store this is unremarkable; it is worth stating only so nobody discovers it as a surprise.

What it deletes

The effort has been looking for what the design removes rather than adds, and this is the first clear instance:

  • Grant kinds. There is one — an exclusive resource. No schema grants, no read roles, no scoping rules for who may see what inside a shared thing.
  • The question of who owns which table, and with it the guessing at revocation time.
  • Cross-module migration ordering. Two modules migrating one database need their migrations ordered against each other. Exclusive ownership means a module's migrations are ordered only against itself.
  • A whole class of permission modelling that a shared store would otherwise need.

The rule is about contexts, not processes

An earlier version of this file argued that a dashboard reading a dozen stores was caught by the rule, because a dashboard is a surface and surfaces speak to an interface. That was wrong, and it drew the line in the wrong place.

The mesh's own board showing nodes, modules and deployments is not a separate context reaching across a boundary. It is the mesh showing its own data. Reading that store is not a violation however it is done, and requiring it to go through an interface to reach facts its own context owns would be ceremony.

The rule is: a context is granted what it exclusively owns. Everything inside that context — its service, its surface, its tools — reads it freely. What is forbidden is a different context reading it.

Which is exactly what the count below shows: the problem was never surfaces. It was three other contexts keeping their tables in the mesh's database.

And one surface over several contexts is normal. The board visualises the mesh, the work engine, the knowledge base and more, and the alternative — a separate web application per context — is worse for everyone who uses it. That is not a compromise with the rule; composing several sources into one view is what a surface is.

What it changes is only where it reads from: each context's interface, not each context's store. Most of that already exists — 56 of 126 modules carry a tool surface, more than carry a service.

And the unified board is what keeps those interfaces honest. If a view cannot be built from a context's interface, that interface is inadequate — discovered in the one place where it is cheap to notice, rather than the first time something else needs the same data and quietly reaches for the store instead.

But "the board calls each context's interface" is not quite it either, and the objection is right: if every context runs its own service with its own interface, the board is coupled to N of them instead of N schemas, something has to compose them, and composition is logic — which tier 3 says a surface does not hold. That moves the problem up a layer rather than solving it.

The skeleton already answers this and the argument above talked past it. work and knowledge are not separate services; they are contexts inside the control plane, alongside the record, inventory, delivery and the rest — and api is listed there as the one interface every surface speaks to.

So the board speaks to one interface. Behind it the contexts stay separate: separate stores, integrating through the record. But they are one tier, one repository, one deployable, and coupling within a tier is not what the tier rule forbids.

Which resolves the objection rather than deflecting it: the problem does move up a layer, and the layer it moves to already exists and has this as its job.

The caveat is load-bearing. This holds only while the contexts are not separate deployables. The moment one becomes its own service with its own interface, the board is back to N clients, something must compose them, and the composition has nowhere to live that tier 3 permits. That is a real constraint on how far the control plane may be split, and it is worth knowing now rather than discovering it by splitting.

If composing even one interface turns out too slow, the answer is a projection the board owns and keeps current from events — not access to somebody else's tables.

Checked against the real consumers

The rule's survival turned on whether every reader of the mesh's own registry could be served some other way. Eighteen consumers open a direct connection to it. Four groups, and only one of them is work.

Group What happens under the rule
The owner and its machinery — the mesh module, the SDK, the environment and configuration synchronisers, secrets Nothing. It owns the database.
Node appliers — the overlay, the shell daemon, the resolver Already resolved. ADR 0016 stops the host querying the mesh database, decided for tier reasons with nothing to do with this.
Foreign tenants — the work engine (10 tables), the knowledge base (2), pipeline logs (1) They need their own database. They are not reading the registry; they are storing their own data in it.
Genuine cross-context reads — the work engine reads nodes; two others read a handful The only ones needing an interface or events.

Thirteen foreign tables live in the mesh's registry database, belonging to three separate contexts. That is how-we-build §4's shared schema, counted.

So the question was the wrong shape. The bulk of the problem is not readers needing a new route to data — it is tenants needing to move out. Tasks, agents and teams have nothing to do with nodes and modules; they are co-located by history. Give that context its own database and its dependency on the registry shrinks to a single table.

What remains is a handful of genuine cross-context reads, small enough to enumerate rather than estimate. The rule holds.

Request or subscription, and what decides

The remaining cross-context reads need one or the other. Neither is SQL — under exclusive ownership a module runs SQL against its own database and nothing else, whatever transport a query might travel over. Both options are the mesh's own channel, and both ride the broker (ADR 0001), so the transport is not the distinction.

The distinction is where the answer lives when you need it.

request subscription
you ask at the moment you need to know never — you are told
the answer lives on the other side in your own store
freshness always current as current as the last event you received
when the other side is down you cannot answer you answer from your copy
what you must handle a round trip that can fail events you missed while you were down

What decides is not taste. ADR 0015 makes disconnection an ordinary situation rather than an exception. So:

Anything that must keep working while disconnected cannot use a request — there is nobody to ask. It needs a local copy, which means a subscription.

And the converse: anything that must be correct at the instant of asking, where a stale answer is worse than no answer, cannot use a subscription. A display can lag. A decision about whether a grant is still valid cannot.

That turns an apparently open question into a derived one. What remains genuinely open is narrower: what a consumer does about the events it missed while it was disconnected — replay from a point, ask once for a full picture and resume, or rebuild from scratch. That is the same question every projection has, and nothing in the record answers it yet.

A migration belongs to the consumer and runs on the provider

A game defines migrations. They run against the database the store granted it. So the migration is:

  • owned by the consumer — it is that module's schema, versioned with that module;
  • hosted by the provider — it runs inside something the consumer does not control;
  • ordered after the provisioning edge — there is nothing to migrate until the grant exists;
  • scoped to the grant — the consumer's migrations touch its database and no other.

Ownership crosses the edge, which nothing in provides and requires expresses. And it is the action category from features.md made concrete: not an artifact, not node state, and not something the host can apply, because the thing it changes is not the machine.

It also gives a consumer's own install an internal order — provisioned, then migrated, then started — that depends on an edge rather than on the module's contents.

What still has no answer

How many instances of postgres should exist? One per mesh is wrong — a node that must work while disconnected cannot depend on a database elsewhere. One per node is wrong — the mesh's own registry is one thing, not one per node. So the answer is per-module, and nothing in the schema says it. This is cases.md axis how many instances, and postgres is the case that proves it cannot be a global rule.

What happens to a grant when the consumer is removed? The game is uninstalled. Its database still exists, holding its data. Dropping it silently is data loss; keeping it forever is a leak. ADR 0016 says the host removes what it applied and no longer declares — but this is not on the host, it is inside another module's state, and the same reasoning does not obviously carry.

Where does node-derived configuration come from? (3) The control plane composes a declaration, and cannot know this machine's memory. Either the host fills in a blank the declaration leaves — which makes the host decide something, against ADR 0016 — or the control plane reads the node's inventory first and composes with it. The second is consistent and means a declaration is composed per node from what the node reported, which is a stronger claim than anything recorded so far.

Providing is not a substrate thing

The four pinned services are the obvious providers, and they are not a category.

Service What a consumer asks it for
a relational store a database, a user, credentials
another relational store, different vendor a database — and not the same one
a message broker a virtual host, a user, permissions
an object store a bucket and keys
an image registry a repository
an identity provider a client, a realm, a secret
an analytics service a site, and a tracking identity
a low-code data platform a base, and a token
an application platform a project, which is several of the above at once
a mail server a mailbox, an alias, credentials

Any hosted service can be a factory. Providing is a facet a module may have, not a kind of module it is — which is the same conclusion the effort reached about services and applications, arriving from the other direction.

That kills the last reason to keep provider as a category. A module runs something, or grants something, or both, or neither.

Two stores, and why database still is not a name

Two relational stores from different vendors both grant a database. They are the sharpest possible test of the substitutability rule from proposal.md, and they fail it completely: different wire protocol, different dialect, different driver, different client library compiled into the consumer.

A consumer declaring requires: database and being handed either would break against one of them. So the name promises what no provider delivers — and now with two real providers in the catalogue rather than a thought experiment.

Database remains a tag. It is how a person finds both. It is not how a consumer names what it needs.

The same shape, three more times

The message broker has all nine. So does the object store, and so does the image registry. They differ in what a consumer asks for — a database, a virtual host, a bucket, a repository — and in nothing structural.

A substrate service is a service plus a factory. That is the whole pattern, and there are four of them. It generalises past the substrate too: anything that grants something per consumer has this shape, and anything that does not is the simpler case.

But two things differ between them, and both matter more than the similarity.

The broker cannot be managed over the broker

ADR 0001 makes the broker the channel every node takes work from, and ADR 0015 makes it the security boundary — everything a node applies arrives through it.

So the module providing the broker is also the way modules are managed. A declaration cannot be delivered to it over itself, and reconfiguring it is done through the thing being reconfigured. Nothing else in the catalogue has that property; the store is consumed by the control plane but is not how the control plane reaches anything.

This is exactly what the carried bundle exists for (ADR 0015): the broker is raised from what the host carries, before there is a channel, because there is no other way to raise it. Recorded here because it is a constraint on one module, not a general rule, and a schema with no way to say so hides it.

Two modules of identical shape want different instance counts

The broker is one per mesh — a single point of failure and a single point of trust, by decision rather than by accident. The store cannot be: a node that must keep working while disconnected (ADR 0015) cannot depend on a database somewhere else.

Same nine properties, opposite answers. Which settles something the cases file left open: how many instances is not derivable from what a module is. It is a decision per module, it has to be declared, and nothing in provides, requires or excludes says it.

And revocation differs in consequence

Revoking a database leaves data behind until something drops it — a leak, and recoverable. Revoking a virtual host drops whatever had not been delivered — not recoverable, and silent.

The relation is the same and the blast radius is not, which is an argument for the provider deciding what revocation means rather than the mesh applying one rule to all of them.

The assignment is a third thing

Recorded because the operator tried the alternative and abandoned it: modules were once node-agnostic, and it did not survive contact.

The worked example says why. Several of a provider's nine properties are not properties of the module at all:

  • where its state lives — a volume on a particular machine;
  • how it is reached — the same module on two nodes may answer locally, on the network, or publicly, and that is a per-assignment decision;
  • configuration derived from the hardware — tuning follows the memory and storage of the machine it landed on;
  • whether this instance is the one a given consumer is provisioned from.

None of those belong in the catalogue, because they differ per node. None belong in the node, because they are about this module. They belong to the pairing, and a design with only modules and nodes has nowhere to put them — which is what "node-agnostic" ran out of.

So there are three entities, not two:

a module · a node · an assignment, which is a module on a node and carries its own configuration

The current system already has this, arrived at the same way: environment values are stored per module and per node, so a module's settings differ between the machines running it.

This does not put placement back in the manifest. A module still says what must be true of a node and never which node (proposal.md). What changes is that the result of placing it is a thing with its own state, rather than a fact recorded on one of the two ends.

And it makes the composed declaration question from above answerable: a declaration is built from the module, the node's inventory, and the assignment between them. Three inputs, which is why two were never enough.