Closes the two open questions in 06 and 08, which turned out to be one question: how many control planes run, and what happens when the hub is down. Both were drifting toward redundancy by default -- a standby plane, a second hub, an election to pick between them. That is not one feature but a property every layer must then honour, and each layer gets it wrong independently. Not wanted, and not needed. A handful of machines with one node hosting the registry is not a distributed system. The argument for why this is sound rather than merely cheap is that the design already tolerates it by construction. ADR 0036 makes reachability state rather than class; the host reconciles from its own store (0043) and never needed to ask anybody to hold the state it was last given. So the control plane being down is not a new failure mode -- it is every node in the ordinary disconnected situation at once. What is lost is change, not operation. No node holds a contended role: the control plane is assigned like any other module, and the overlay hub is declared (0050). No promotion, no quorum, no fencing, no split brain, no replicated store, and no "which node is authoritative" recurring at every layer. Two consequences stated plainly rather than buried. The control-plane node is a single point of failure -- deliberate, and said out loud so it stays deliberate. And recovery is restore rather than failover, which makes backup the availability story rather than hygiene. The sharpest one is the clock: the control plane owns certificate issuance (0049), so an outage outlasting a renewal window expires every public name. That bounds how long recovery may take, and nothing measures it today.
03-DESIGN
The authoritative specification. Implementation is built against what is written here.
Two layers
| Folder | What it is |
|---|---|
00-as-is/ |
The mesh that exists today. Shipped behaviour, described as it is — including behaviour nobody would choose again. |
01-to-be/ |
The mesh being built toward. Every statement traceable to a record in 02-DECISIONS/. |
They are never mixed. A statement about the future does not belong in an as-is document, and an as-is document is never edited to describe an intention.
When a to-be design ships, it does not move. Its as-is counterpart is written or updated,
the to-be document's status becomes implemented, and both stand — one describing what runs,
the other recording what was intended. Deleting the intention loses the reasoning, which is
the expensive half.
Frontmatter
Every design document (not the READMEs) carries:
---
layer: as-is | to-be
status: designed | in-progress | implemented | abandoned
code: [] # owning code repo(s), from 00-META/repos.md
updated: YYYY-MM-DD # date of the last status change, not of text edits
decisions: [] # 02-DECISIONS/ records this document rests on
---
For an as-is document, status: implemented is the normal state — it describes something that
runs — and code: names where that implementation lives.
Status changes when implementation state changes, never because design text was edited. An
implemented claim must be defensible from the owning repository's main branch, not from
intent. If it cannot be checked, it is in-progress.
Cross-cutting views are generated from this frontmatter by the hq-status skill and never
written to disk.
What belongs here
Functional analysis, architectural description, and specification — prose and diagrams
only, no code. A manifest field may be named; a manifest may not be pasted. A document
enters the to-be layer only after the decision behind it is recorded in 02-DECISIONS/
and the research that produced it is closed.
Subfolders are encouraged where a layer grows enough to need them.