The 'domain grouping' item cited ADR 0017 as live guidance. 0044 superseded it -- there is no domain module to group into, so there is no domain list to settle.
6.4 KiB
status, date, deciders, reconstructed, extends
| status | date | deciders | reconstructed | extends |
|---|---|---|---|---|
| accepted | 2026-08-27 | jochen | false | 0036-a-node-is-a-managed-machine.md |
53. One node runs the control plane, and nothing takes over
Context
Two design documents left the same question open from opposite ends:
06-the-control-plane.md— how many nodes run it, and what happens when the one running it is down.08-connectivity.md— what happens when the hub is down, given that WireGuard has no failover and every non-co-located pair routes through it.
Both were drifting toward the same answer by default: redundancy. A second hub, a standby control plane, an election to decide which is live. That direction is expensive in a specific way — it is not one feature but a property that every layer must then honour, and each layer gets it wrong independently.
It is also not wanted. This is a mesh of a handful of machines with one node hosting the registry, not a distributed system, and building for a failure mode nobody asked to survive would buy nothing while complicating everything.
Considered options
- A standby control plane with promotion. Needs a replicated store, an election, a fencing mechanism so two promoted planes cannot both write, and a rehearsed promotion procedure — which is only trustworthy if it is practised, and an unpractised failover is reliably worse than none. Rejected.
- Multiple control planes, each authoritative for part of the mesh. Trades availability for a partition problem and a merge problem, and contradicts ADR 0045's exclusive ownership at the worst possible layer. Rejected.
- One, declared, with no failover. Chosen.
Decision
One node runs the control plane and hosts its store. It is declared, never elected, and nothing takes over when it is down.
No node holds a contended role. A node runs the control plane because it was assigned to,
by the same mechanism that assigns anything else
(ADR 0002, 06).
There is no promotion, no election, no quorum, no consensus, and therefore no split brain. The
overlay hub is declared the same way and for the same reason
(ADR 0050).
Why this is sound, and not merely cheap
The design already tolerates the control plane being absent, and it tolerates it by construction rather than by luck.
ADR 0036 settles that reachability is state, not class — a disconnected node has a last-known state and a pending set of declarations. The host reconciles from its own store (ADR 0043), not from the mesh, and ADR 0037 means it never needed to ask anybody in order to keep a machine in the state it was last told to hold.
So:
The control plane being down is not a new failure mode. It is every node in the ordinary disconnected situation, at the same time.
That is the argument. A property the design already has for one node does not stop being true because it applies to all of them at once. What is lost is change, not operation — every node goes on running exactly what it was last told to run.
What is actually lost while it is down
Worth listing, because "nodes keep working" is true and is not the whole picture:
| new assignments, new modules, new versions | stop |
| new grants and provisioning | stop |
| a new node joining | stops — the token is issued by the mesh |
| collection of health, logs and reports | stops |
| non-co-located overlay paths | stop — the hub is the route |
| co-located direct peers | keep working |
| everything already assigned, on every node | keeps running |
Consequences
- A large amount of machinery is never built, and this is the point: leader election, quorum, fencing, a replicated store, split-brain reconciliation, promotion runbooks, and the question which node is authoritative recurring at every layer. None of it exists, so none of it can be subtly wrong.
- The control-plane node is a single point of failure, and the design says so plainly. That is a deliberate position, not an oversight, and stating it is what keeps it deliberate — an undocumented single point of failure is discovered during the outage.
- Recovery is restore, not failover — so backup becomes the availability story. It stops being hygiene and becomes the mechanism the whole arrangement rests on. An unverified backup here is not a risk to the backup; it is the mesh having no recovery path at all. Research 012 is where that lives, and it is now load-bearing rather than prudent.
- Certificate renewal is the clock, and it is the sharpest consequence. Nodes keep running indefinitely, but the control plane owns issuance (ADR 0049), so an outage lasting longer than a renewal window expires public certificates and takes down every public name. That converts an inconvenience into an outage on a timer, and it is the real bound on how long recovery may take — not patience, not the number of stopped features.
- The bound is not measured anywhere. Nothing today reports how close a certificate is to expiry or how long the control plane has been unreachable, and both are needed for the above to be a plan rather than a hope.
- It can be revisited without unpicking anything. Nothing here assumes singularity in a way that would have to be undone — the store is exclusively owned, the host is independent, and the hub is declared. Adding redundancy later means adding it, not reversing this.