Files
hq/03-DESIGN/01-to-be/06-the-control-plane.md
T
jschoubben ccbbfa9c8a One node runs the control plane, and nothing takes over
Closes the two open questions in 06 and 08, which turned out to be one
question: how many control planes run, and what happens when the hub is down.
Both were drifting toward redundancy by default -- a standby plane, a second
hub, an election to pick between them. That is not one feature but a property
every layer must then honour, and each layer gets it wrong independently.

Not wanted, and not needed. A handful of machines with one node hosting the
registry is not a distributed system.

The argument for why this is sound rather than merely cheap is that the design
already tolerates it by construction. ADR 0036 makes reachability state rather
than class; the host reconciles from its own store (0043) and never needed to
ask anybody to hold the state it was last given. So the control plane being
down is not a new failure mode -- it is every node in the ordinary disconnected
situation at once. What is lost is change, not operation.

No node holds a contended role: the control plane is assigned like any other
module, and the overlay hub is declared (0050). No promotion, no quorum, no
fencing, no split brain, no replicated store, and no "which node is
authoritative" recurring at every layer.

Two consequences stated plainly rather than buried. The control-plane node is a
single point of failure -- deliberate, and said out loud so it stays
deliberate. And recovery is restore rather than failover, which makes backup
the availability story rather than hygiene.

The sharpest one is the clock: the control plane owns certificate issuance
(0049), so an outage outlasting a renewal window expires every public name.
That bounds how long recovery may take, and nothing measures it today.
2026-08-27 00:55:10 +02:00

135 lines
6.8 KiB
Markdown

---
layer: to-be
status: designed
code: []
updated: 2026-08-27
decisions:
- 02-DECISIONS/0037-the-host-applies-it-does-not-decide.md
- 02-DECISIONS/0053-one-control-plane-and-no-failover.md
- 02-DECISIONS/0030-the-repository-structure.md
- 02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md
---
# The control plane
Tier 2. The term appears seventy-nine times across this repository and was defined nowhere,
which is `how-we-build` §5 failing on this repository's own vocabulary.
This document defines it. It does **not** design the contexts inside it; those are open in
[research 006](../../01-RESEARCH/006-mesh-from-scratch/00-overview.md).
## The definition
> **The control plane is everything that needs to know about more than one node.**
That is the whole test, and it is not arbitrary — it follows from
[ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md). The host applies and
does not decide *because deciding needs knowledge the machine does not have*. So the line falls
exactly there:
| Question | Whose |
|---|---|
| write this file, with this content, with this mode | the **host** — one machine |
| which nodes should run the store | the **control plane** — needs every node |
| is this unit running | the **host** — one machine |
| which peers belong in this node's overlay | the **control plane** — needs every node |
| what does this machine have installed | the **host** reports; the control plane **records** |
| has this node been unreachable for a week | the **control plane** — nobody else is watching |
A useful consequence: **anything a single machine could answer alone is not the control
plane's.** If it needs no second node, putting it here is a mistake, and the tier rule will not
catch it because the dependency direction is still correct.
## What is inside it
Ten contexts and one interface, from the skeleton
([research 006](../../01-RESEARCH/006-mesh-from-scratch/skeleton.md)):
| | |
|---|---|
| **record** | the event log every other context integrates through |
| **inventory** | nodes, modules, assignments, versions |
| **config** | settings, secrets, and deriving them onto nodes |
| **connectivity** | overlay, resolution, exposure, filtering, certificates — **specified in full in [`08-connectivity.md`](08-connectivity.md)** |
| **provisioning** | resource grants between modules |
| **delivery** | source to artifact to node |
| **observability** | health, logs, metrics, alerts |
| **identity** | agents, humans, services, authorisation |
| **work** | tasks, workflows, runs |
| **knowledge** | memory, documents, retrieval |
| **api** | the one interface every surface speaks to |
**These are contexts, not services.** They are separate in the sense that matters — each owns
its own store, and they integrate through the record rather than by reading one another
([`how-we-build`](../../00-META/how-we-build.md) §4). They are not separate deployables, and
[research 011](../../01-RESEARCH/011-the-module-graph/00-overview.md) records why that
constraint is load-bearing: a single surface can compose them only while there is one interface
in front of them.
## What it is not
- **Not the thing that changes machines.** It decides; the host applies. It never reaches into a
node except through the host.
- **Not a surface.** Tier 3 is how people and agents reach it. It has one interface; the
surfaces are what speak to that interface.
- **Not the substrate.** It *runs on* tier 1 — PostgreSQL, LavinMQ, MinIO, an OCI registry
([ADR 0048](../../02-DECISIONS/0048-the-substrate-is-named.md)) — and cannot start without
them, which is what makes them a lower tier.
- **Not privileged on a node.** It has no more access to a machine than the declaration
vocabulary allows ([ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md)).
## It is also a consumer
The property that makes tier 2 unlike the others: **the control plane has requirements of its
own.** It needs a PostgreSQL database, an AMQP virtual host, and a bucket — the same things any
module needs, granted the same way.
That is the circularity the tiers exist to resolve rather than hide: the control plane cannot
provision its own database, because it is not running yet. So its **store** is raised from the
bundle the host carries, before there is a control plane to ask
([ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md),
[research 011](../../01-RESEARCH/011-the-module-graph/worked-provider.md)).
Its virtual host and its bucket are **not** in the bundle — by the time they are wanted there is
a control plane to grant them. Whether the bus must come first is
[open](07-the-substrate.md#open), and it turns on whether these contexts talk to each other over
it.
## Where it runs
**On nodes, like anything else.** It is not a place outside the mesh; it is modules the mesh
hosts, assigned to nodes by the same mechanism as everything else.
**One node runs it, and nothing takes over**
([ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md)). The node is assigned,
never elected — no promotion, no quorum, no split brain.
That is sound rather than merely cheap, because the design already tolerates the control plane
being absent by construction: a node reconciles from **its own** store
([ADR 0043](../../02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md)) and
never needed to ask anybody to hold the state it was last given. So the control plane being down
is not a new failure mode — it is
[ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)'s ordinary disconnected
situation, happening to every node at once. **What is lost is change, not operation.**
The honest half: this node is a single point of failure, recovery is restore rather than
failover, and **certificate renewal is the clock** — an outage outlasting a renewal window expires
every public name.
## Open
- **The contexts themselves.** Ten is the skeleton's claim, not a settled list. Research 006
asks whether the record belongs here or in the substrate, and whether identity is a context or
a substrate service.
- **How far it may be split.** One deployable today. Splitting a context out costs the single
interface a surface depends on
([research 011](../../01-RESEARCH/011-the-module-graph/00-overview.md)).
- ~~**How many run, and what a node does without one.**~~ **Resolved** by
[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md). What remains is
measurement: nothing reports how long the control plane has been unreachable, or how close a
certificate is to expiry — both needed for restore-not-failover to be a plan rather than a
hope.
- **What the interface is.** One interface is stated; its shape, and whether it is request,
subscription or both, is not
([research 011](../../01-RESEARCH/011-the-module-graph/worked-provider.md)).