Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.
the node host 8 -> 1 applies not decides, depends on nothing,
per operating system, root service, the
launcher, episodic, what a declaration is,
actions from the bundle only
a node and how it joins 4 -> 1 what a node is, joining, the link as
security boundary, the enrolment token
modules and the graph 7 -> 1 everything is a module, no domain modules,
three edges, provisioning, the core library
substrate and control 6 -> 1 the test, seven contexts, one control plane,
plane the authority is not a database, the named
products, the pinned bundle
connectivity 3 -> 1 a route is a grant, reachability declared,
filter rules
delivery 5 -> 1 reconciliation not a pipeline, artifacts,
the three silos, a failed step, the verdict
the lab 5 -> 1 (earlier)
how this repository 10 -> 1 (earlier)
works
Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.
The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.
The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
13 KiB
layer, status, code, updated, decisions
| layer | status | code | updated | decisions | |||||
|---|---|---|---|---|---|---|---|---|---|
| to-be | designed | 2026-08-27 |
|
The control plane
Tier 2. The term appears seventy-nine times across this repository and was defined nowhere,
which is how-we-build §5 failing on this repository's own vocabulary.
This document defines it. It does not design the contexts inside it; those are open in research 006.
The definition
The control plane is everything that needs to know about more than one node.
That is the whole test, and it is not arbitrary — it follows from ADR 0037. The host applies and does not decide because deciding needs knowledge the machine does not have. So the line falls exactly there:
| Question | Whose |
|---|---|
| write this file, with this content, with this mode | the host — one machine |
| which nodes should run the store | the control plane — needs every node |
| is this unit running | the host — one machine |
| which peers belong in this node's overlay | the control plane — needs every node |
| what does this machine have installed | the host reports; the control plane records |
| has this node been unreachable for a week | the control plane — nobody else is watching |
A useful consequence: anything a single machine could answer alone is not the control plane's. If it needs no second node, putting it here is a mistake, and the tier rule will not catch it because the dependency direction is still correct.
What is inside it
Seven contexts and one interface (ADR 0048) — each one earning its place by the test above rather than by being ours:
| needs to know about more than one node because | ||
|---|---|---|
| inventory | nodes, modules, assignments, versions | that is the mesh-wide fact |
| config | settings, secrets, and deriving them onto nodes | it derives onto nodes |
| connectivity | overlay, resolution, exposure, filtering, certificates — specified in full in 08-connectivity.md |
who peers with whom; which node is reachable |
| provisioning | resource grants between modules | consumer and provider may be on different nodes |
| delivery | source to artifact to node | it targets nodes |
| observability | health, logs, metrics, alerts | unreachable for a week is nobody else's to notice |
| identity | agents, humans, services, authorisation | credentials follow an agent's node bindings and modality |
| api | the one interface every surface speaks to | — it is an interface, not a context |
What is deliberately not here. work, knowledge and stream are mesh-hosted
applications — first-party, shipped with everything else, and running on the mesh the way
anything else does. A task does not need to know a node exists, and being ours does not make
something infrastructure. ai is folded into config: a provider licence is an ordinary grant.
record is an open question rather than an eighth entry. Contexts integrate through it
(ADR 0045), which makes it load-bearing,
and research 006 leaves where it lives
unresolved — putting it in the substrate risks recreating the circularity the tier design just
removed. Listing it here would settle by naming what has not been settled by arguing.
These are contexts, not services. They are separate in the sense that matters — each owns
its own store, and they integrate through the record rather than by reading one another
(how-we-build §4). They are not separate deployables, and
research 011 records why that
constraint is load-bearing: a single surface can compose them only while there is one interface
in front of them.
Nothing outside a context touches its store
The question this answers: can a node write to the registry database? No — and not "only through one node", which is the weaker arrangement it might be mistaken for.
No node holds a credential to any control-plane store, for writing or for reading.
That is not a new rule here; it is four already taken, and it is worth seeing them together because each one alone reads like a detail:
| ADR 0037 | the host never queries the mesh database |
| ADR 0036 | a node holds its own identity and nothing else — the shared database credential every node carries today is the exposure this exists to remove |
| ADR 0045 | a context is granted only what it exclusively owns: no shared writes, no read-only roles |
| ADR 0048 | there is no single mesh database, and nothing reads one |
So how does anything get in
Over the broker, as a message; the owning context writes.
node ──event/report──► broker ──► the context that owns that data ──► its own store
A node reports what it applied, what it holds, and that it is alive. It states; it does not write. The difference is the whole security boundary: a node that can write cannot be prevented from writing anything, and a node that can only state has its blast radius bounded by what the message vocabulary can say.
Reads work the same way in reverse — a node is told, in declarations. It never asks.
Who actually consumes, and who writes
The control plane is the consumer. There is one of it, and the context that owns the data does the write.
node ──► broker ──► the control plane, consuming
├─ a node reported what it applied ─► inventory writes the registry
├─ a node reported health ─► observability writes its own store
└─ a grant was requested ─► provisioning writes its own store
Seven contexts, one deployable — they are not separate services, so this is one process consuming and dispatching internally, not seven consumers racing. Each context then writes only the store it exclusively owns (ADR 0045).
One consumer is a property worth having, not just a consequence of ADR 0048. The as-is records that two consumers accidentally sharing one queue silently split the traffic between them, each receiving half of what it expects — which has happened, between a module's daemon and its capability server. With one consumer that class of fault cannot arise.
And the broker is the buffer while the control plane is down. Nodes go on publishing; messages queue; the control plane drains them when it returns. That is what makes ADR 0048's single control plane tolerable — an outage delays the mesh's knowledge rather than losing it.
With one consequence that must be bounded before it is discovered: a queue with no limit grows until the broker's disk is full, and the broker is the one component every node depends on. Queues carrying node reports need a maximum length or a message lifetime, and losing the oldest health report is obviously right where losing the oldest declaration acknowledgement is not — so the bound is per queue and is not decided here.
On volume, which is the real worry underneath
Most high-frequency writes are not registry writes, and that is the first thing to check
before designing for throughput. The registry is inventory's store: nodes, modules,
assignments, versions. Those change when somebody changes something.
Logs, metrics and health checks belong to observability, which owns a different store
(ADR 0045). Sending them to the registry
would be exactly the shared-schema mistake 0045 exists to stop, arriving through the back door
marked performance.
That leaves one genuine funnel: every context's writes go through the process that owns it, and ADR 0048 says there is one of it. For a mesh of a handful of machines this is not a scaling problem, and it should not be solved by giving nodes database credentials — that trades a bounded problem for an unbounded one. If it ever binds, the answers are at the consumer: batch, apply backpressure, or move the highest volume stream out of a relational store entirely.
What observability actually stores its data in is not decided, and it is the one place where volume genuinely argues against a relational store.
What it is not
- Not the thing that changes machines. It decides; the host applies. It never reaches into a node except through the host.
- Not a surface. Tier 3 is how people and agents reach it. It has one interface; the surfaces are what speak to that interface.
- Not the substrate. It runs on tier 1 — PostgreSQL, LavinMQ, MinIO, an OCI registry (ADR 0048) — and cannot start without them, which is what makes them a lower tier.
- Not privileged on a node. It has no more access to a machine than the declaration vocabulary allows (ADR 0036).
It is also a consumer
The property that makes tier 2 unlike the others: the control plane has requirements of its own. It needs a PostgreSQL database, an AMQP virtual host, and a bucket — the same things any module needs, granted the same way.
That is the circularity the tiers exist to resolve rather than hide: the control plane cannot provision its own database, because it is not running yet. So its store is raised from the bundle the host carries, before there is a control plane to ask (ADR 0036, research 011).
Its virtual host and its bucket are not in the bundle — by the time they are wanted there is a control plane to grant them. Whether the bus must come first is open, and it turns on whether these contexts talk to each other over it.
Where it runs
On nodes, like anything else. It is not a place outside the mesh; it is modules the mesh hosts, assigned to nodes by the same mechanism as everything else.
One node runs it, and nothing takes over (ADR 0048). The node is assigned, never elected — no promotion, no quorum, no split brain.
That is sound rather than merely cheap, because the design already tolerates the control plane being absent by construction: a node reconciles from its own store (ADR 0037) and never needed to ask anybody to hold the state it was last given. So the control plane being down is not a new failure mode — it is ADR 0036's ordinary disconnected situation, happening to every node at once. What is lost is change, not operation.
The honest half: this node is a single point of failure, recovery is restore rather than failover, and certificate renewal is the clock — an outage outlasting a renewal window expires every public name.
Open
The contexts themselves.Decided — seven, by ADR 0048. What remains open is narrower and named there: where the record lives, which research 006 leaves unresolved because the substrate is the one place it must not go.- How far it may be split. One deployable today. Splitting a context out costs the single interface a surface depends on (research 011).
How many run, and what a node does without one.Resolved by ADR 0048. What remains is measurement: nothing reports how long the control plane has been unreachable, or how close a certificate is to expiry — both needed for restore-not-failover to be a plan rather than a hope.- What the interface is. One interface is stated; its shape, and whether it is request, subscription or both, is not (research 011).