One node runs the control plane, and nothing takes over

Closes the two open questions in 06 and 08, which turned out to be one
question: how many control planes run, and what happens when the hub is down.
Both were drifting toward redundancy by default -- a standby plane, a second
hub, an election to pick between them. That is not one feature but a property
every layer must then honour, and each layer gets it wrong independently.

Not wanted, and not needed. A handful of machines with one node hosting the
registry is not a distributed system.

The argument for why this is sound rather than merely cheap is that the design
already tolerates it by construction. ADR 0036 makes reachability state rather
than class; the host reconciles from its own store (0043) and never needed to
ask anybody to hold the state it was last given. So the control plane being
down is not a new failure mode -- it is every node in the ordinary disconnected
situation at once. What is lost is change, not operation.

No node holds a contended role: the control plane is assigned like any other
module, and the overlay hub is declared (0050). No promotion, no quorum, no
fencing, no split brain, no replicated store, and no "which node is
authoritative" recurring at every layer.

Two consequences stated plainly rather than buried. The control-plane node is a
single point of failure -- deliberate, and said out loud so it stays
deliberate. And recovery is restore rather than failover, which makes backup
the availability story rather than hygiene.

The sharpest one is the clock: the control plane owns certificate issuance
(0049), so an outage outlasting a renewal window expires every public name.
That bounds how long recovery may take, and nothing measures it today.
This commit is contained in:
2026-08-27 00:55:10 +02:00
parent 4e80820e2f
commit ccbbfa9c8a
3 changed files with 151 additions and 11 deletions
+21 -6
View File
@@ -5,6 +5,7 @@ code: []
updated: 2026-08-27
decisions:
- 02-DECISIONS/0037-the-host-applies-it-does-not-decide.md
- 02-DECISIONS/0053-one-control-plane-and-no-failover.md
- 02-DECISIONS/0030-the-repository-structure.md
- 02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md
---
@@ -99,11 +100,21 @@ it.
**On nodes, like anything else.** It is not a place outside the mesh; it is modules the mesh
hosts, assigned to nodes by the same mechanism as everything else.
Which raises a question this document does not answer: **how many nodes run it, and what happens
when the one running it is down.** The broker is one per mesh by decision; whether the control
plane is, and what a node does while it cannot reach it, is
[ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)'s ordinary situation seen from
the other end — and it is not designed.
**One node runs it, and nothing takes over**
([ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md)). The node is assigned,
never elected — no promotion, no quorum, no split brain.
That is sound rather than merely cheap, because the design already tolerates the control plane
being absent by construction: a node reconciles from **its own** store
([ADR 0043](../../02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md)) and
never needed to ask anybody to hold the state it was last given. So the control plane being down
is not a new failure mode — it is
[ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)'s ordinary disconnected
situation, happening to every node at once. **What is lost is change, not operation.**
The honest half: this node is a single point of failure, recovery is restore rather than
failover, and **certificate renewal is the clock** — an outage outlasting a renewal window expires
every public name.
## Open
@@ -113,7 +124,11 @@ the other end — and it is not designed.
- **How far it may be split.** One deployable today. Splitting a context out costs the single
interface a surface depends on
([research 011](../../01-RESEARCH/011-the-module-graph/00-overview.md)).
- **How many run, and what a node does without one.** Above.
- ~~**How many run, and what a node does without one.**~~ **Resolved** by
[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md). What remains is
measurement: nothing reports how long the control plane has been unreachable, or how close a
certificate is to expiry — both needed for restore-not-failover to be a plan rather than a
hope.
- **What the interface is.** One interface is stated; its shape, and whether it is request,
subscription or both, is not
([research 011](../../01-RESEARCH/011-the-module-graph/worked-provider.md)).
+7 -5
View File
@@ -10,6 +10,7 @@ decisions:
- 02-DECISIONS/0050-reachability-is-a-property-of-the-address.md
- 02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md
- 02-DECISIONS/0052-a-filter-rule-names-its-source.md
- 02-DECISIONS/0053-one-control-plane-and-no-failover.md
---
# Connectivity
@@ -217,11 +218,12 @@ The list is worth having in one place, because it is most of the argument:
## Open
- **What happens when the hub is down.** WireGuard has no failover, and the hub is a single point
through which every non-co-located pair routes. This design does not add a second hub and does
not pretend the first one is redundant. It is the same shape as
[`06`](06-the-control-plane.md)'s *how many run, and what a node does without one*, and it
should be answered with it rather than separately.
- ~~**What happens when the hub is down.**~~ **Resolved** by
[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md), together with `06`'s
matching question — they were one question. Nothing takes over. WireGuard has no failover, the
hub is declared rather than elected, and non-co-located paths stop while co-located direct peers
and every already-assigned workload keep running. The recovery path is restore, and its deadline
is certificate renewal.
- **Renumbering the overlay.** Made *possible* by declaring the hub rather than inferring it from
an address, but no procedure exists, and a graph delivered node by node has an ordering problem
while it is half-applied.