Closes the two open questions in 06 and 08, which turned out to be one question: how many control planes run, and what happens when the hub is down. Both were drifting toward redundancy by default -- a standby plane, a second hub, an election to pick between them. That is not one feature but a property every layer must then honour, and each layer gets it wrong independently. Not wanted, and not needed. A handful of machines with one node hosting the registry is not a distributed system. The argument for why this is sound rather than merely cheap is that the design already tolerates it by construction. ADR 0036 makes reachability state rather than class; the host reconciles from its own store (0043) and never needed to ask anybody to hold the state it was last given. So the control plane being down is not a new failure mode -- it is every node in the ordinary disconnected situation at once. What is lost is change, not operation. No node holds a contended role: the control plane is assigned like any other module, and the overlay hub is declared (0050). No promotion, no quorum, no fencing, no split brain, no replicated store, and no "which node is authoritative" recurring at every layer. Two consequences stated plainly rather than buried. The control-plane node is a single point of failure -- deliberate, and said out loud so it stays deliberate. And recovery is restore rather than failover, which makes backup the availability story rather than hygiene. The sharpest one is the clock: the control plane owns certificate issuance (0049), so an outage outlasting a renewal window expires every public name. That bounds how long recovery may take, and nothing measures it today.
03-DESIGN / 01-to-be
The mesh being built toward. Every statement here traces to a record in
02-DECISIONS/; nothing arrives by drafting.
A document here describes an intention. What currently runs is in
00-as-is/, and the two are never merged — when something ships, the as-is
document is written and this one's status becomes implemented.
| Document | Covers | Rests on |
|---|---|---|
00-work-breakdown.md |
How the decomposition gets built, in what order, and where a human must look | ADR 0015 |
01-end-to-end-testing.md |
The lab: a real mesh a change can be run against before it reaches nodes | ADR 0016, 0029 |
02-scenario-declaration.md |
What a scenario declares — the underlay, and what to place on it | ADR 0031 |
03-scenario-lifecycle.md |
What happens to a scenario — raise, snapshot, restore, move, destroy | ADR 0032 |
04-lab-installation.md |
Getting the lab onto a clean machine, and why it verifies capability rather than installation | ADR 0008 |
05-the-node-host.md |
Tier 0 — the one thing installed by hand, and the only thing that changes a machine | ADR 0037 |
06-the-control-plane.md |
Tier 2 — what the term means, and the test for what belongs in it | ADR 0037 |
07-the-substrate.md |
Tier 1 — what the control plane consumes and cannot grant itself | ADR 0038, 0048 |
08-connectivity.md |
One context in full — overlay, resolution, exposure, filtering, certificates | ADR 0049, 0050, 0051 |
Not yet written
- The remaining contexts. ADR 0015
decides the decomposition;
connectivityis the first written in full (08) and the others do not exist yet. The work breakdown says in what order they are needed. - Domain grouping outside the core. ADR 0017 settles the principle and explicitly does not settle the domain list. That is a research effort, not a design document, until it concludes.