Files
hq/03-DESIGN/01-to-be/06-the-control-plane.md
T
jschoubben 19997d56c3 Approve 0057-0059, with four corrections from review
Not approved as drafted -- four things came out of checking them against each
other, and one was a bug that would have broken every upgrade.

The bug: 0059 specified Restart=on-failure while 0057 has the host restart onto
a new binary by exiting CLEANLY. on-failure does not restart a process that
exited zero, so every upgraded node would have been left stopped, having
successfully upgraded. Found by reading the two records against each other
rather than by either alone. Now Restart=always in all three places that
mention it.

The host cannot run in a container, and the reason is decisive rather than
stylistic: step 0 of the substrate bootstrap installs the container runtime, so
a host inside a container would need the thing it exists to install. It would
also break 0041 -- copy it onto a machine and run it stops being true when the
machine must already have a runtime. Everything above tier 0 is a container;
the host is not. That split is the tier boundary, not an inconsistency.

systemd is named rather than abstracted. An init is not a dependency in 0041's
sense: 0041 is about what must be installed before the host works, and an init
is not installed, it is what the machine already is. The unit file is the only
systemd-specific artefact and it belongs to the package, so a machine with a
different supervisor ships a different package.

The mesh is a watchdog, and my first draft was half an answer. Recovery must be
local -- nothing dials a node, and a host that cannot start cannot report. But
detection is the mesh's, and a local supervisor structurally cannot do it: it
sees one process failing and cannot tell a broken machine from a broken
release. Only something watching every node can, and that distinction decides
whether the response is "fix this machine" or "stop shipping this version". So
a host rollout is staged -- a few nodes, wait for heartbeats, continue or stop
on silence. Local rollback still needed, because the canary nodes break and
because a node offline during the rollout gets the declaration later with no
batch around it.

The first declaration is the overlay and nothing else. Forced, because a node's
address and peers are assigned rather than chosen. But also the way back in: a
node reachable over the overlay can be fixed by hand if a later declaration
breaks it, and a large first declaration risks a node that is broken and
unreachable at once.

Also stated plainly, because it reads as a contradiction: nodes reach each
other over the overlay and every node consumes from the broker; what 0039
forbids is an inbound CONTROL surface, not reachability.

And in 06: no node holds a credential to any control-plane store, for reads or
writes. Four ADRs already say this separately and none of them said it in one
place. Nodes state over the broker; the owning context writes. With a note that
most high-frequency writes are observability's, not the registry's -- routing
logs into the registry would be the shared-schema mistake arriving through a
door marked performance.
2026-08-27 22:12:12 +02:00

199 lines
11 KiB
Markdown

---
layer: to-be
status: designed
code: []
updated: 2026-08-27
decisions:
- 02-DECISIONS/0037-the-host-applies-it-does-not-decide.md
- 02-DECISIONS/0053-one-control-plane-and-no-failover.md
- 02-DECISIONS/0030-the-repository-structure.md
- 02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md
- 02-DECISIONS/0055-the-control-plane-is-the-node-coordinating-contexts.md
---
# The control plane
Tier 2. The term appears seventy-nine times across this repository and was defined nowhere,
which is `how-we-build` §5 failing on this repository's own vocabulary.
This document defines it. It does **not** design the contexts inside it; those are open in
[research 006](../../01-RESEARCH/006-mesh-from-scratch/00-overview.md).
## The definition
> **The control plane is everything that needs to know about more than one node.**
That is the whole test, and it is not arbitrary — it follows from
[ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md). The host applies and
does not decide *because deciding needs knowledge the machine does not have*. So the line falls
exactly there:
| Question | Whose |
|---|---|
| write this file, with this content, with this mode | the **host** — one machine |
| which nodes should run the store | the **control plane** — needs every node |
| is this unit running | the **host** — one machine |
| which peers belong in this node's overlay | the **control plane** — needs every node |
| what does this machine have installed | the **host** reports; the control plane **records** |
| has this node been unreachable for a week | the **control plane** — nobody else is watching |
A useful consequence: **anything a single machine could answer alone is not the control
plane's.** If it needs no second node, putting it here is a mistake, and the tier rule will not
catch it because the dependency direction is still correct.
## What is inside it
**Seven contexts and one interface**
([ADR 0055](../../02-DECISIONS/0055-the-control-plane-is-the-node-coordinating-contexts.md)) —
each one earning its place by the test above rather than by being ours:
| | | needs to know about more than one node because |
|---|---|---|
| **inventory** | nodes, modules, assignments, versions | that *is* the mesh-wide fact |
| **config** | settings, secrets, and deriving them onto nodes | it derives **onto nodes** |
| **connectivity** | overlay, resolution, exposure, filtering, certificates — **specified in full in [`08-connectivity.md`](08-connectivity.md)** | who peers with whom; which node is reachable |
| **provisioning** | resource grants between modules | consumer and provider may be on different nodes |
| **delivery** | source to artifact to node | it targets nodes |
| **observability** | health, logs, metrics, alerts | *unreachable for a week* is nobody else's to notice |
| **identity** | agents, humans, services, authorisation | credentials follow an agent's node bindings and modality |
| **api** | the one interface every surface speaks to | — it is an interface, not a context |
**What is deliberately not here.** `work`, `knowledge` and `stream` are **mesh-hosted
applications** — first-party, shipped with everything else, and running on the mesh the way
anything else does. A task does not need to know a node exists, and *being ours does not make
something infrastructure*. `ai` is folded into `config`: a provider licence is an ordinary grant.
**`record` is an open question rather than an eighth entry.** Contexts integrate through it
([ADR 0045](../../02-DECISIONS/0045-a-context-owns-its-store.md)), which makes it load-bearing,
and [research 006](../../01-RESEARCH/006-mesh-from-scratch/skeleton.md) leaves *where it lives*
unresolved — putting it in the substrate risks recreating the circularity the tier design just
removed. Listing it here would settle by naming what has not been settled by arguing.
**These are contexts, not services.** They are separate in the sense that matters — each owns
its own store, and they integrate through the record rather than by reading one another
([`how-we-build`](../../00-META/how-we-build.md) §4). They are not separate deployables, and
[research 011](../../01-RESEARCH/011-the-module-graph/00-overview.md) records why that
constraint is load-bearing: a single surface can compose them only while there is one interface
in front of them.
## Nothing outside a context touches its store
The question this answers: **can a node write to the registry database?** No — and not "only
through one node", which is the weaker arrangement it might be mistaken for.
> **No node holds a credential to any control-plane store, for writing or for reading.**
That is not a new rule here; it is four already taken, and it is worth seeing them together
because each one alone reads like a detail:
| | |
|---|---|
| [ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md) | the host never queries the mesh database |
| [ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md) | a node holds its own identity **and nothing else** — the shared database credential every node carries today is the exposure this exists to remove |
| [ADR 0045](../../02-DECISIONS/0045-a-context-owns-its-store.md) | a context is granted only what it **exclusively** owns: no shared writes, no read-only roles |
| [ADR 0056](../../02-DECISIONS/0056-the-authority-is-the-control-plane-not-a-database.md) | there is no single mesh database, and nothing reads one |
### So how does anything get in
**Over the broker, as a message; the owning context writes.**
```
node ──event/report──► broker ──► the context that owns that data ──► its own store
```
A node reports what it applied, what it holds, and that it is alive. It **states**; it does not
**write**. The difference is the whole security boundary: a node that can write cannot be
prevented from writing anything, and a node that can only state has its blast radius bounded by
what the message vocabulary can say.
Reads work the same way in reverse — a node is *told*, in declarations. It never asks.
### On volume, which is the real worry underneath
**Most high-frequency writes are not registry writes, and that is the first thing to check
before designing for throughput.** The registry is `inventory`'s store: nodes, modules,
assignments, versions. Those change when somebody changes something.
**Logs, metrics and health checks belong to `observability`**, which owns a different store
([ADR 0045](../../02-DECISIONS/0045-a-context-owns-its-store.md)). Sending them to the registry
would be exactly the shared-schema mistake 0045 exists to stop, arriving through the back door
marked *performance*.
That leaves one genuine funnel: every context's writes go through the process that owns it, and
[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md) says there is one of it.
For a mesh of a handful of machines this is not a scaling problem, and **it should not be solved
by giving nodes database credentials** — that trades a bounded problem for an unbounded one. If
it ever binds, the answers are at the consumer: batch, apply backpressure, or move the highest
volume stream out of a relational store entirely.
**What observability actually stores its data in is not decided**, and it is the one place where
volume genuinely argues against a relational store.
## What it is not
- **Not the thing that changes machines.** It decides; the host applies. It never reaches into a
node except through the host.
- **Not a surface.** Tier 3 is how people and agents reach it. It has one interface; the
surfaces are what speak to that interface.
- **Not the substrate.** It *runs on* tier 1 — PostgreSQL, LavinMQ, MinIO, an OCI registry
([ADR 0048](../../02-DECISIONS/0048-the-substrate-is-named.md)) — and cannot start without
them, which is what makes them a lower tier.
- **Not privileged on a node.** It has no more access to a machine than the declaration
vocabulary allows ([ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md)).
## It is also a consumer
The property that makes tier 2 unlike the others: **the control plane has requirements of its
own.** It needs a PostgreSQL database, an AMQP virtual host, and a bucket — the same things any
module needs, granted the same way.
That is the circularity the tiers exist to resolve rather than hide: the control plane cannot
provision its own database, because it is not running yet. So its **store** is raised from the
bundle the host carries, before there is a control plane to ask
([ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md),
[research 011](../../01-RESEARCH/011-the-module-graph/worked-provider.md)).
Its virtual host and its bucket are **not** in the bundle — by the time they are wanted there is
a control plane to grant them. Whether the bus must come first is
[open](07-the-substrate.md#open), and it turns on whether these contexts talk to each other over
it.
## Where it runs
**On nodes, like anything else.** It is not a place outside the mesh; it is modules the mesh
hosts, assigned to nodes by the same mechanism as everything else.
**One node runs it, and nothing takes over**
([ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md)). The node is assigned,
never elected — no promotion, no quorum, no split brain.
That is sound rather than merely cheap, because the design already tolerates the control plane
being absent by construction: a node reconciles from **its own** store
([ADR 0043](../../02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md)) and
never needed to ask anybody to hold the state it was last given. So the control plane being down
is not a new failure mode — it is
[ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)'s ordinary disconnected
situation, happening to every node at once. **What is lost is change, not operation.**
The honest half: this node is a single point of failure, recovery is restore rather than
failover, and **certificate renewal is the clock** — an outage outlasting a renewal window expires
every public name.
## Open
- ~~**The contexts themselves.**~~ **Decided** — seven, by
[ADR 0055](../../02-DECISIONS/0055-the-control-plane-is-the-node-coordinating-contexts.md).
What remains open is narrower and named there: **where the record lives**, which research 006
leaves unresolved because the substrate is the one place it must not go.
- **How far it may be split.** One deployable today. Splitting a context out costs the single
interface a surface depends on
([research 011](../../01-RESEARCH/011-the-module-graph/00-overview.md)).
- ~~**How many run, and what a node does without one.**~~ **Resolved** by
[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md). What remains is
measurement: nothing reports how long the control plane has been unreachable, or how close a
certificate is to expiry — both needed for restore-not-failover to be a plan rather than a
hope.
- **What the interface is.** One interface is stated; its shape, and whether it is request,
subscription or both, is not
([research 011](../../01-RESEARCH/011-the-module-graph/worked-provider.md)).