One node runs the control plane, and nothing takes over

Closes the two open questions in 06 and 08, which turned out to be one
question: how many control planes run, and what happens when the hub is down.
Both were drifting toward redundancy by default -- a standby plane, a second
hub, an election to pick between them. That is not one feature but a property
every layer must then honour, and each layer gets it wrong independently.

Not wanted, and not needed. A handful of machines with one node hosting the
registry is not a distributed system.

The argument for why this is sound rather than merely cheap is that the design
already tolerates it by construction. ADR 0036 makes reachability state rather
than class; the host reconciles from its own store (0043) and never needed to
ask anybody to hold the state it was last given. So the control plane being
down is not a new failure mode -- it is every node in the ordinary disconnected
situation at once. What is lost is change, not operation.

No node holds a contended role: the control plane is assigned like any other
module, and the overlay hub is declared (0050). No promotion, no quorum, no
fencing, no split brain, no replicated store, and no "which node is
authoritative" recurring at every layer.

Two consequences stated plainly rather than buried. The control-plane node is a
single point of failure -- deliberate, and said out loud so it stays
deliberate. And recovery is restore rather than failover, which makes backup
the availability story rather than hygiene.

The sharpest one is the clock: the control plane owns certificate issuance
(0049), so an outage outlasting a renewal window expires every public name.
That bounds how long recovery may take, and nothing measures it today.
This commit is contained in:
2026-08-27 00:55:10 +02:00
parent 4e80820e2f
commit ccbbfa9c8a
3 changed files with 151 additions and 11 deletions
@@ -0,0 +1,123 @@
---
status: proposed
date: 2026-08-27
deciders: jochen
reconstructed: false
extends: 0036-a-node-is-a-managed-machine.md
---
# 53. One node runs the control plane, and nothing takes over
## Context
Two design documents left the same question open from opposite ends:
- [`06-the-control-plane.md`](../03-DESIGN/01-to-be/06-the-control-plane.md) — *how many nodes
run it, and what happens when the one running it is down.*
- [`08-connectivity.md`](../03-DESIGN/01-to-be/08-connectivity.md) — *what happens when the hub
is down*, given that WireGuard has no failover and every non-co-located pair routes through it.
Both were drifting toward the same answer by default: redundancy. A second hub, a standby control
plane, an election to decide which is live. That direction is expensive in a specific way — it is
not one feature but a property that every layer must then honour, and each layer gets it wrong
independently.
**It is also not wanted.** This is a mesh of a handful of machines with one node hosting the
registry, not a distributed system, and building for a failure mode nobody asked to survive would
buy nothing while complicating everything.
## Considered options
1. **A standby control plane with promotion.** Needs a replicated store, an election, a fencing
mechanism so two promoted planes cannot both write, and a rehearsed promotion procedure —
which is only trustworthy if it is *practised*, and an unpractised failover is reliably worse
than none. Rejected.
2. **Multiple control planes, each authoritative for part of the mesh.** Trades availability for
a partition problem and a merge problem, and contradicts
[ADR 0045](0045-a-context-owns-its-store.md)'s exclusive ownership at the worst possible layer.
Rejected.
3. **One, declared, with no failover.** Chosen.
## Decision
**One node runs the control plane and hosts its store. It is declared, never elected, and nothing
takes over when it is down.**
**No node holds a contended role.** A node runs the control plane because it was *assigned* to,
by the same mechanism that assigns anything else
([ADR 0002](0002-everything-is-a-module.md), [`06`](../03-DESIGN/01-to-be/06-the-control-plane.md)).
There is no promotion, no election, no quorum, no consensus, and therefore no split brain. The
overlay hub is declared the same way and for the same reason
([ADR 0050](0050-reachability-is-a-property-of-the-address.md)).
### Why this is sound, and not merely cheap
**The design already tolerates the control plane being absent, and it tolerates it by
construction rather than by luck.**
[ADR 0036](0036-a-node-is-a-managed-machine.md) settles that *reachability is state, not class* —
a disconnected node has a last-known state and a pending set of declarations. The host reconciles
from **its own** store ([ADR 0043](0043-a-declaration-is-an-ordered-list-of-owned-resources.md)),
not from the mesh, and [ADR 0037](0037-the-host-applies-it-does-not-decide.md) means it never
needed to ask anybody in order to keep a machine in the state it was last told to hold.
So:
> **The control plane being down is not a new failure mode. It is every node in the ordinary
> disconnected situation, at the same time.**
That is the argument. A property the design already has for one node does not stop being true
because it applies to all of them at once. **What is lost is change, not operation** — every node
goes on running exactly what it was last told to run.
### What is actually lost while it is down
Worth listing, because "nodes keep working" is true and is not the whole picture:
| | |
|---|---|
| new assignments, new modules, new versions | **stop** |
| new grants and provisioning | **stop** |
| a new node joining | **stops** — the token is issued by the mesh |
| collection of health, logs and reports | **stops** |
| non-co-located overlay paths | **stop** — the hub is the route |
| co-located direct peers | keep working |
| everything already assigned, on every node | **keeps running** |
## Consequences
- **A large amount of machinery is never built**, and this is the point: leader election, quorum,
fencing, a replicated store, split-brain reconciliation, promotion runbooks, and the question
*which node is authoritative* recurring at every layer. None of it exists, so none of it can be
subtly wrong.
- **The control-plane node is a single point of failure, and the design says so plainly.** That is
a deliberate position, not an oversight, and stating it is what keeps it deliberate — an
undocumented single point of failure is discovered during the outage.
- **Recovery is restore, not failover — so backup becomes the availability story.** It stops being
hygiene and becomes the mechanism the whole arrangement rests on. An unverified backup here is
not a risk to the backup; it is the mesh having no recovery path at all.
[Research 012](../01-RESEARCH/012-the-minimum-viable-node/00-overview.md) is where that lives,
and it is now load-bearing rather than prudent.
- **Certificate renewal is the clock, and it is the sharpest consequence.** Nodes keep running
indefinitely, but the control plane owns issuance
([ADR 0049](0049-a-route-is-a-grant.md)), so an outage lasting longer than a renewal window
expires public certificates and takes down every public name. **That converts an inconvenience
into an outage on a timer**, and it is the real bound on how long recovery may take — not
patience, not the number of stopped features.
- **The bound is not measured anywhere.** Nothing today reports how close a certificate is to
expiry or how long the control plane has been unreachable, and both are needed for the above to
be a plan rather than a hope.
- **It can be revisited without unpicking anything.** Nothing here assumes singularity in a way
that would have to be undone — the store is exclusively owned, the host is independent, and the
hub is declared. Adding redundancy later means adding it, not reversing this.
## References
- [ADR 0036](0036-a-node-is-a-managed-machine.md) — disconnection as an ordinary situation, which
is the whole argument.
- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) and
[ADR 0043](0043-a-declaration-is-an-ordered-list-of-owned-resources.md) — why a node keeps
running without anybody to ask.
- [ADR 0049](0049-a-route-is-a-grant.md) — certificate issuance, which sets the recovery clock.
- [`06`](../03-DESIGN/01-to-be/06-the-control-plane.md), [`08`](../03-DESIGN/01-to-be/08-connectivity.md) —
the two open questions this closes.
+21 -6
View File
@@ -5,6 +5,7 @@ code: []
updated: 2026-08-27
decisions:
- 02-DECISIONS/0037-the-host-applies-it-does-not-decide.md
- 02-DECISIONS/0053-one-control-plane-and-no-failover.md
- 02-DECISIONS/0030-the-repository-structure.md
- 02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md
---
@@ -99,11 +100,21 @@ it.
**On nodes, like anything else.** It is not a place outside the mesh; it is modules the mesh
hosts, assigned to nodes by the same mechanism as everything else.
Which raises a question this document does not answer: **how many nodes run it, and what happens
when the one running it is down.** The broker is one per mesh by decision; whether the control
plane is, and what a node does while it cannot reach it, is
[ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)'s ordinary situation seen from
the other end — and it is not designed.
**One node runs it, and nothing takes over**
([ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md)). The node is assigned,
never elected — no promotion, no quorum, no split brain.
That is sound rather than merely cheap, because the design already tolerates the control plane
being absent by construction: a node reconciles from **its own** store
([ADR 0043](../../02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md)) and
never needed to ask anybody to hold the state it was last given. So the control plane being down
is not a new failure mode — it is
[ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)'s ordinary disconnected
situation, happening to every node at once. **What is lost is change, not operation.**
The honest half: this node is a single point of failure, recovery is restore rather than
failover, and **certificate renewal is the clock** — an outage outlasting a renewal window expires
every public name.
## Open
@@ -113,7 +124,11 @@ the other end — and it is not designed.
- **How far it may be split.** One deployable today. Splitting a context out costs the single
interface a surface depends on
([research 011](../../01-RESEARCH/011-the-module-graph/00-overview.md)).
- **How many run, and what a node does without one.** Above.
- ~~**How many run, and what a node does without one.**~~ **Resolved** by
[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md). What remains is
measurement: nothing reports how long the control plane has been unreachable, or how close a
certificate is to expiry — both needed for restore-not-failover to be a plan rather than a
hope.
- **What the interface is.** One interface is stated; its shape, and whether it is request,
subscription or both, is not
([research 011](../../01-RESEARCH/011-the-module-graph/worked-provider.md)).
+7 -5
View File
@@ -10,6 +10,7 @@ decisions:
- 02-DECISIONS/0050-reachability-is-a-property-of-the-address.md
- 02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md
- 02-DECISIONS/0052-a-filter-rule-names-its-source.md
- 02-DECISIONS/0053-one-control-plane-and-no-failover.md
---
# Connectivity
@@ -217,11 +218,12 @@ The list is worth having in one place, because it is most of the argument:
## Open
- **What happens when the hub is down.** WireGuard has no failover, and the hub is a single point
through which every non-co-located pair routes. This design does not add a second hub and does
not pretend the first one is redundant. It is the same shape as
[`06`](06-the-control-plane.md)'s *how many run, and what a node does without one*, and it
should be answered with it rather than separately.
- ~~**What happens when the hub is down.**~~ **Resolved** by
[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md), together with `06`'s
matching question — they were one question. Nothing takes over. WireGuard has no failover, the
hub is declared rather than elected, and non-co-located paths stop while co-located direct peers
and every already-assigned workload keep running. The recovery path is restore, and its deadline
is certificate renewal.
- **Renumbering the overlay.** Made *possible* by declaring the hub rather than inferring it from
an address, but no procedure exists, and a graph delivered node by node has an ordering problem
while it is half-applied.