--- status: accepted date: 2026-08-27 deciders: jochen reconstructed: false extends: 0036-a-node-is-a-managed-machine.md --- # 53. One node runs the control plane, and nothing takes over ## Context Two design documents left the same question open from opposite ends: - [`06-the-control-plane.md`](../03-DESIGN/01-to-be/06-the-control-plane.md) — *how many nodes run it, and what happens when the one running it is down.* - [`08-connectivity.md`](../03-DESIGN/01-to-be/08-connectivity.md) — *what happens when the hub is down*, given that WireGuard has no failover and every non-co-located pair routes through it. Both were drifting toward the same answer by default: redundancy. A second hub, a standby control plane, an election to decide which is live. That direction is expensive in a specific way — it is not one feature but a property that every layer must then honour, and each layer gets it wrong independently. **It is also not wanted.** This is a mesh of a handful of machines with one node hosting the registry, not a distributed system, and building for a failure mode nobody asked to survive would buy nothing while complicating everything. ## Considered options 1. **A standby control plane with promotion.** Needs a replicated store, an election, a fencing mechanism so two promoted planes cannot both write, and a rehearsed promotion procedure — which is only trustworthy if it is *practised*, and an unpractised failover is reliably worse than none. Rejected. 2. **Multiple control planes, each authoritative for part of the mesh.** Trades availability for a partition problem and a merge problem, and contradicts [ADR 0045](0045-a-context-owns-its-store.md)'s exclusive ownership at the worst possible layer. Rejected. 3. **One, declared, with no failover.** Chosen. ## Decision **One node runs the control plane and hosts its store. It is declared, never elected, and nothing takes over when it is down.** **No node holds a contended role.** A node runs the control plane because it was *assigned* to, by the same mechanism that assigns anything else ([ADR 0002](0002-everything-is-a-module.md), [`06`](../03-DESIGN/01-to-be/06-the-control-plane.md)). There is no promotion, no election, no quorum, no consensus, and therefore no split brain. The overlay hub is declared the same way and for the same reason ([ADR 0050](0050-reachability-is-a-property-of-the-address.md)). ### Why this is sound, and not merely cheap **The design already tolerates the control plane being absent, and it tolerates it by construction rather than by luck.** [ADR 0036](0036-a-node-is-a-managed-machine.md) settles that *reachability is state, not class* — a disconnected node has a last-known state and a pending set of declarations. The host reconciles from **its own** store ([ADR 0043](0043-a-declaration-is-an-ordered-list-of-owned-resources.md)), not from the mesh, and [ADR 0037](0037-the-host-applies-it-does-not-decide.md) means it never needed to ask anybody in order to keep a machine in the state it was last told to hold. So: > **The control plane being down is not a new failure mode. It is every node in the ordinary > disconnected situation, at the same time.** That is the argument. A property the design already has for one node does not stop being true because it applies to all of them at once. **What is lost is change, not operation** — every node goes on running exactly what it was last told to run. ### What is actually lost while it is down Worth listing, because "nodes keep working" is true and is not the whole picture: | | | |---|---| | new assignments, new modules, new versions | **stop** | | new grants and provisioning | **stop** | | a new node joining | **stops** — the token is issued by the mesh | | collection of health, logs and reports | **stops** | | non-co-located overlay paths | **stop** — the hub is the route | | co-located direct peers | keep working | | everything already assigned, on every node | **keeps running** | ## Consequences - **A large amount of machinery is never built**, and this is the point: leader election, quorum, fencing, a replicated store, split-brain reconciliation, promotion runbooks, and the question *which node is authoritative* recurring at every layer. None of it exists, so none of it can be subtly wrong. - **The control-plane node is a single point of failure, and the design says so plainly.** That is a deliberate position, not an oversight, and stating it is what keeps it deliberate — an undocumented single point of failure is discovered during the outage. - **Recovery is restore, not failover — so backup becomes the availability story.** It stops being hygiene and becomes the mechanism the whole arrangement rests on. An unverified backup here is not a risk to the backup; it is the mesh having no recovery path at all. [Research 012](../01-RESEARCH/012-the-minimum-viable-node/00-overview.md) is where that lives, and it is now load-bearing rather than prudent. - **Certificate renewal is the clock, and it is the sharpest consequence.** Nodes keep running indefinitely, but the control plane owns issuance ([ADR 0049](0049-a-route-is-a-grant.md)), so an outage lasting longer than a renewal window expires public certificates and takes down every public name. **That converts an inconvenience into an outage on a timer**, and it is the real bound on how long recovery may take — not patience, not the number of stopped features. - **The bound is not measured anywhere.** Nothing today reports how close a certificate is to expiry or how long the control plane has been unreachable, and both are needed for the above to be a plan rather than a hope. - **It can be revisited without unpicking anything.** Nothing here assumes singularity in a way that would have to be undone — the store is exclusively owned, the host is independent, and the hub is declared. Adding redundancy later means adding it, not reversing this. ## References - [ADR 0036](0036-a-node-is-a-managed-machine.md) — disconnection as an ordinary situation, which is the whole argument. - [ADR 0037](0037-the-host-applies-it-does-not-decide.md) and [ADR 0043](0043-a-declaration-is-an-ordered-list-of-owned-resources.md) — why a node keeps running without anybody to ask. - [ADR 0049](0049-a-route-is-a-grant.md) — certificate issuance, which sets the recovery clock. - [`06`](../03-DESIGN/01-to-be/06-the-control-plane.md), [`08`](../03-DESIGN/01-to-be/08-connectivity.md) — the two open questions this closes.