Files
hq/02-DECISIONS/0053-one-control-plane-and-no-failover.md
T
jschoubben ef5dd0751b Approve 0049-0053; drop a to-be item superseded by ADR 0044
The 'domain grouping' item cited ADR 0017 as live guidance. 0044 superseded
it -- there is no domain module to group into, so there is no domain list to
settle.
2026-08-27 01:00:29 +02:00

6.4 KiB

status, date, deciders, reconstructed, extends
status date deciders reconstructed extends
accepted 2026-08-27 jochen false 0036-a-node-is-a-managed-machine.md

53. One node runs the control plane, and nothing takes over

Context

Two design documents left the same question open from opposite ends:

  • 06-the-control-plane.md — how many nodes run it, and what happens when the one running it is down.
  • 08-connectivity.md — what happens when the hub is down, given that WireGuard has no failover and every non-co-located pair routes through it.

Both were drifting toward the same answer by default: redundancy. A second hub, a standby control plane, an election to decide which is live. That direction is expensive in a specific way — it is not one feature but a property that every layer must then honour, and each layer gets it wrong independently.

It is also not wanted. This is a mesh of a handful of machines with one node hosting the registry, not a distributed system, and building for a failure mode nobody asked to survive would buy nothing while complicating everything.

Considered options

  1. A standby control plane with promotion. Needs a replicated store, an election, a fencing mechanism so two promoted planes cannot both write, and a rehearsed promotion procedure — which is only trustworthy if it is practised, and an unpractised failover is reliably worse than none. Rejected.
  2. Multiple control planes, each authoritative for part of the mesh. Trades availability for a partition problem and a merge problem, and contradicts ADR 0045's exclusive ownership at the worst possible layer. Rejected.
  3. One, declared, with no failover. Chosen.

Decision

One node runs the control plane and hosts its store. It is declared, never elected, and nothing takes over when it is down.

No node holds a contended role. A node runs the control plane because it was assigned to, by the same mechanism that assigns anything else (ADR 0002, 06). There is no promotion, no election, no quorum, no consensus, and therefore no split brain. The overlay hub is declared the same way and for the same reason (ADR 0050).

Why this is sound, and not merely cheap

The design already tolerates the control plane being absent, and it tolerates it by construction rather than by luck.

ADR 0036 settles that reachability is state, not class — a disconnected node has a last-known state and a pending set of declarations. The host reconciles from its own store (ADR 0043), not from the mesh, and ADR 0037 means it never needed to ask anybody in order to keep a machine in the state it was last told to hold.

So:

The control plane being down is not a new failure mode. It is every node in the ordinary disconnected situation, at the same time.

That is the argument. A property the design already has for one node does not stop being true because it applies to all of them at once. What is lost is change, not operation — every node goes on running exactly what it was last told to run.

What is actually lost while it is down

Worth listing, because "nodes keep working" is true and is not the whole picture:

new assignments, new modules, new versions stop
new grants and provisioning stop
a new node joining stops — the token is issued by the mesh
collection of health, logs and reports stops
non-co-located overlay paths stop — the hub is the route
co-located direct peers keep working
everything already assigned, on every node keeps running

Consequences

  • A large amount of machinery is never built, and this is the point: leader election, quorum, fencing, a replicated store, split-brain reconciliation, promotion runbooks, and the question which node is authoritative recurring at every layer. None of it exists, so none of it can be subtly wrong.
  • The control-plane node is a single point of failure, and the design says so plainly. That is a deliberate position, not an oversight, and stating it is what keeps it deliberate — an undocumented single point of failure is discovered during the outage.
  • Recovery is restore, not failover — so backup becomes the availability story. It stops being hygiene and becomes the mechanism the whole arrangement rests on. An unverified backup here is not a risk to the backup; it is the mesh having no recovery path at all. Research 012 is where that lives, and it is now load-bearing rather than prudent.
  • Certificate renewal is the clock, and it is the sharpest consequence. Nodes keep running indefinitely, but the control plane owns issuance (ADR 0049), so an outage lasting longer than a renewal window expires public certificates and takes down every public name. That converts an inconvenience into an outage on a timer, and it is the real bound on how long recovery may take — not patience, not the number of stopped features.
  • The bound is not measured anywhere. Nothing today reports how close a certificate is to expiry or how long the control plane has been unreachable, and both are needed for the above to be a plan rather than a hope.
  • It can be revisited without unpicking anything. Nothing here assumes singularity in a way that would have to be undone — the store is exclusively owned, the host is independent, and the hub is declared. Adding redundancy later means adding it, not reversing this.

References

  • ADR 0036 — disconnection as an ordinary situation, which is the whole argument.
  • ADR 0037 and ADR 0043 — why a node keeps running without anybody to ask.
  • ADR 0049 — certificate issuance, which sets the recovery clock.
  • 06, 08 — the two open questions this closes.