One node runs the control plane, and nothing takes over
Closes the two open questions in 06 and 08, which turned out to be one question: how many control planes run, and what happens when the hub is down. Both were drifting toward redundancy by default -- a standby plane, a second hub, an election to pick between them. That is not one feature but a property every layer must then honour, and each layer gets it wrong independently. Not wanted, and not needed. A handful of machines with one node hosting the registry is not a distributed system. The argument for why this is sound rather than merely cheap is that the design already tolerates it by construction. ADR 0036 makes reachability state rather than class; the host reconciles from its own store (0043) and never needed to ask anybody to hold the state it was last given. So the control plane being down is not a new failure mode -- it is every node in the ordinary disconnected situation at once. What is lost is change, not operation. No node holds a contended role: the control plane is assigned like any other module, and the overlay hub is declared (0050). No promotion, no quorum, no fencing, no split brain, no replicated store, and no "which node is authoritative" recurring at every layer. Two consequences stated plainly rather than buried. The control-plane node is a single point of failure -- deliberate, and said out loud so it stays deliberate. And recovery is restore rather than failover, which makes backup the availability story rather than hygiene. The sharpest one is the clock: the control plane owns certificate issuance (0049), so an outage outlasting a renewal window expires every public name. That bounds how long recovery may take, and nothing measures it today.
This commit is contained in:
@@ -5,6 +5,7 @@ code: []
|
||||
updated: 2026-08-27
|
||||
decisions:
|
||||
- 02-DECISIONS/0037-the-host-applies-it-does-not-decide.md
|
||||
- 02-DECISIONS/0053-one-control-plane-and-no-failover.md
|
||||
- 02-DECISIONS/0030-the-repository-structure.md
|
||||
- 02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md
|
||||
---
|
||||
@@ -99,11 +100,21 @@ it.
|
||||
**On nodes, like anything else.** It is not a place outside the mesh; it is modules the mesh
|
||||
hosts, assigned to nodes by the same mechanism as everything else.
|
||||
|
||||
Which raises a question this document does not answer: **how many nodes run it, and what happens
|
||||
when the one running it is down.** The broker is one per mesh by decision; whether the control
|
||||
plane is, and what a node does while it cannot reach it, is
|
||||
[ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)'s ordinary situation seen from
|
||||
the other end — and it is not designed.
|
||||
**One node runs it, and nothing takes over**
|
||||
([ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md)). The node is assigned,
|
||||
never elected — no promotion, no quorum, no split brain.
|
||||
|
||||
That is sound rather than merely cheap, because the design already tolerates the control plane
|
||||
being absent by construction: a node reconciles from **its own** store
|
||||
([ADR 0043](../../02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md)) and
|
||||
never needed to ask anybody to hold the state it was last given. So the control plane being down
|
||||
is not a new failure mode — it is
|
||||
[ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)'s ordinary disconnected
|
||||
situation, happening to every node at once. **What is lost is change, not operation.**
|
||||
|
||||
The honest half: this node is a single point of failure, recovery is restore rather than
|
||||
failover, and **certificate renewal is the clock** — an outage outlasting a renewal window expires
|
||||
every public name.
|
||||
|
||||
## Open
|
||||
|
||||
@@ -113,7 +124,11 @@ the other end — and it is not designed.
|
||||
- **How far it may be split.** One deployable today. Splitting a context out costs the single
|
||||
interface a surface depends on
|
||||
([research 011](../../01-RESEARCH/011-the-module-graph/00-overview.md)).
|
||||
- **How many run, and what a node does without one.** Above.
|
||||
- ~~**How many run, and what a node does without one.**~~ **Resolved** by
|
||||
[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md). What remains is
|
||||
measurement: nothing reports how long the control plane has been unreachable, or how close a
|
||||
certificate is to expiry — both needed for restore-not-failover to be a plan rather than a
|
||||
hope.
|
||||
- **What the interface is.** One interface is stated; its shape, and whether it is request,
|
||||
subscription or both, is not
|
||||
([research 011](../../01-RESEARCH/011-the-module-graph/worked-provider.md)).
|
||||
|
||||
@@ -10,6 +10,7 @@ decisions:
|
||||
- 02-DECISIONS/0050-reachability-is-a-property-of-the-address.md
|
||||
- 02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md
|
||||
- 02-DECISIONS/0052-a-filter-rule-names-its-source.md
|
||||
- 02-DECISIONS/0053-one-control-plane-and-no-failover.md
|
||||
---
|
||||
|
||||
# Connectivity
|
||||
@@ -217,11 +218,12 @@ The list is worth having in one place, because it is most of the argument:
|
||||
|
||||
## Open
|
||||
|
||||
- **What happens when the hub is down.** WireGuard has no failover, and the hub is a single point
|
||||
through which every non-co-located pair routes. This design does not add a second hub and does
|
||||
not pretend the first one is redundant. It is the same shape as
|
||||
[`06`](06-the-control-plane.md)'s *how many run, and what a node does without one*, and it
|
||||
should be answered with it rather than separately.
|
||||
- ~~**What happens when the hub is down.**~~ **Resolved** by
|
||||
[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md), together with `06`'s
|
||||
matching question — they were one question. Nothing takes over. WireGuard has no failover, the
|
||||
hub is declared rather than elected, and non-co-located paths stop while co-located direct peers
|
||||
and every already-assigned workload keep running. The recovery path is restore, and its deadline
|
||||
is certificate renewal.
|
||||
- **Renumbering the overlay.** Made *possible* by declaring the hub rather than inferring it from
|
||||
an address, but no procedure exists, and a graph delivered node by node has an ordering problem
|
||||
while it is half-applied.
|
||||
|
||||
Reference in New Issue
Block a user