Files
hq/03-DESIGN
jschoubben ccbbfa9c8a One node runs the control plane, and nothing takes over
Closes the two open questions in 06 and 08, which turned out to be one
question: how many control planes run, and what happens when the hub is down.
Both were drifting toward redundancy by default -- a standby plane, a second
hub, an election to pick between them. That is not one feature but a property
every layer must then honour, and each layer gets it wrong independently.

Not wanted, and not needed. A handful of machines with one node hosting the
registry is not a distributed system.

The argument for why this is sound rather than merely cheap is that the design
already tolerates it by construction. ADR 0036 makes reachability state rather
than class; the host reconciles from its own store (0043) and never needed to
ask anybody to hold the state it was last given. So the control plane being
down is not a new failure mode -- it is every node in the ordinary disconnected
situation at once. What is lost is change, not operation.

No node holds a contended role: the control plane is assigned like any other
module, and the overlay hub is declared (0050). No promotion, no quorum, no
fencing, no split brain, no replicated store, and no "which node is
authoritative" recurring at every layer.

Two consequences stated plainly rather than buried. The control-plane node is a
single point of failure -- deliberate, and said out loud so it stays
deliberate. And recovery is restore rather than failover, which makes backup
the availability story rather than hygiene.

The sharpest one is the clock: the control plane owns certificate issuance
(0049), so an outage outlasting a renewal window expires every public name.
That bounds how long recovery may take, and nothing measures it today.
2026-08-27 00:55:10 +02:00
..
2026-08-25 01:55:52 +02:00

03-DESIGN

The authoritative specification. Implementation is built against what is written here.

Two layers

Folder What it is
00-as-is/ The mesh that exists today. Shipped behaviour, described as it is — including behaviour nobody would choose again.
01-to-be/ The mesh being built toward. Every statement traceable to a record in 02-DECISIONS/.

They are never mixed. A statement about the future does not belong in an as-is document, and an as-is document is never edited to describe an intention.

When a to-be design ships, it does not move. Its as-is counterpart is written or updated, the to-be document's status becomes implemented, and both stand — one describing what runs, the other recording what was intended. Deleting the intention loses the reasoning, which is the expensive half.

Frontmatter

Every design document (not the READMEs) carries:

---
layer: as-is | to-be
status: designed | in-progress | implemented | abandoned
code: []                 # owning code repo(s), from 00-META/repos.md
updated: YYYY-MM-DD      # date of the last status change, not of text edits
decisions: []            # 02-DECISIONS/ records this document rests on
---

For an as-is document, status: implemented is the normal state — it describes something that runs — and code: names where that implementation lives.

Status changes when implementation state changes, never because design text was edited. An implemented claim must be defensible from the owning repository's main branch, not from intent. If it cannot be checked, it is in-progress.

Cross-cutting views are generated from this frontmatter by the hq-status skill and never written to disk.

What belongs here

Functional analysis, architectural description, and specification — prose and diagrams only, no code. A manifest field may be named; a manifest may not be pasted. A document enters the to-be layer only after the decision behind it is recorded in 02-DECISIONS/ and the research that produced it is closed.

Subfolders are encouraged where a layer grows enough to need them.