One node runs the control plane, and nothing takes over

Closes the two open questions in 06 and 08, which turned out to be one
question: how many control planes run, and what happens when the hub is down.
Both were drifting toward redundancy by default -- a standby plane, a second
hub, an election to pick between them. That is not one feature but a property
every layer must then honour, and each layer gets it wrong independently.

Not wanted, and not needed. A handful of machines with one node hosting the
registry is not a distributed system.

The argument for why this is sound rather than merely cheap is that the design
already tolerates it by construction. ADR 0036 makes reachability state rather
than class; the host reconciles from its own store (0043) and never needed to
ask anybody to hold the state it was last given. So the control plane being
down is not a new failure mode -- it is every node in the ordinary disconnected
situation at once. What is lost is change, not operation.

No node holds a contended role: the control plane is assigned like any other
module, and the overlay hub is declared (0050). No promotion, no quorum, no
fencing, no split brain, no replicated store, and no "which node is
authoritative" recurring at every layer.

Two consequences stated plainly rather than buried. The control-plane node is a
single point of failure -- deliberate, and said out loud so it stays
deliberate. And recovery is restore rather than failover, which makes backup
the availability story rather than hygiene.

The sharpest one is the clock: the control plane owns certificate issuance
(0049), so an outage outlasting a renewal window expires every public name.
That bounds how long recovery may take, and nothing measures it today.
This commit is contained in:
2026-08-27 00:55:10 +02:00
parent 4e80820e2f
commit ccbbfa9c8a
3 changed files with 151 additions and 11 deletions
+7 -5
View File
@@ -10,6 +10,7 @@ decisions:
- 02-DECISIONS/0050-reachability-is-a-property-of-the-address.md
- 02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md
- 02-DECISIONS/0052-a-filter-rule-names-its-source.md
- 02-DECISIONS/0053-one-control-plane-and-no-failover.md
---
# Connectivity
@@ -217,11 +218,12 @@ The list is worth having in one place, because it is most of the argument:
## Open
- **What happens when the hub is down.** WireGuard has no failover, and the hub is a single point
through which every non-co-located pair routes. This design does not add a second hub and does
not pretend the first one is redundant. It is the same shape as
[`06`](06-the-control-plane.md)'s *how many run, and what a node does without one*, and it
should be answered with it rather than separately.
- ~~**What happens when the hub is down.**~~ **Resolved** by
[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md), together with `06`'s
matching question — they were one question. Nothing takes over. WireGuard has no failover, the
hub is declared rather than elected, and non-co-located paths stop while co-located direct peers
and every already-assigned workload keep running. The recovery path is restore, and its deadline
is certificate renewal.
- **Renumbering the overlay.** Made *possible* by declaring the hub rather than inferring it from
an address, but no procedure exists, and a graph delivered node by node has an ordering problem
while it is half-applied.