Files
hq/03-DESIGN/01-to-be/08-connectivity.md
T
jschoubben ccbbfa9c8a One node runs the control plane, and nothing takes over
Closes the two open questions in 06 and 08, which turned out to be one
question: how many control planes run, and what happens when the hub is down.
Both were drifting toward redundancy by default -- a standby plane, a second
hub, an election to pick between them. That is not one feature but a property
every layer must then honour, and each layer gets it wrong independently.

Not wanted, and not needed. A handful of machines with one node hosting the
registry is not a distributed system.

The argument for why this is sound rather than merely cheap is that the design
already tolerates it by construction. ADR 0036 makes reachability state rather
than class; the host reconciles from its own store (0043) and never needed to
ask anybody to hold the state it was last given. So the control plane being
down is not a new failure mode -- it is every node in the ordinary disconnected
situation at once. What is lost is change, not operation.

No node holds a contended role: the control plane is assigned like any other
module, and the overlay hub is declared (0050). No promotion, no quorum, no
fencing, no split brain, no replicated store, and no "which node is
authoritative" recurring at every layer.

Two consequences stated plainly rather than buried. The control-plane node is a
single point of failure -- deliberate, and said out loud so it stays
deliberate. And recovery is restore rather than failover, which makes backup
the availability story rather than hygiene.

The sharpest one is the clock: the control plane owns certificate issuance
(0049), so an outage outlasting a renewal window expires every public name.
That bounds how long recovery may take, and nothing measures it today.
2026-08-27 00:55:10 +02:00

12 KiB

layer, status, code, updated, decisions
layer status code updated decisions
to-be designed
2026-08-27
02-DECISIONS/0037-the-host-applies-it-does-not-decide.md
02-DECISIONS/0039-the-link-is-the-security-boundary.md
02-DECISIONS/0049-a-route-is-a-grant.md
02-DECISIONS/0050-reachability-is-a-property-of-the-address.md
02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md
02-DECISIONS/0052-a-filter-rule-names-its-source.md
02-DECISIONS/0053-one-control-plane-and-no-failover.md

Connectivity

One of the control plane's ten contexts, and the one with the most moving parts: overlay, resolution, exposure, filtering, certificates.

It is written as a whole because the five are one design. They share inputs, they must agree, and every one of them today is computed in a different place by a different module from a different copy of the same facts.

Why it is control-plane work

Apply the test — everything that needs to know about more than one node — to each responsibility:

needs to know whose
overlay — who peers with whom, at what address every node, and which of them can be dialled control plane
resolution — which name is which node every node control plane
exposure — which public name reaches which container which node is publicly reachable (ADR 0049) control plane
filtering — which port is open, to whom what is assigned here, and the overlay's shape control plane decides, host applies
certificates — who may present which name which name belongs to which node control plane

Not one of the five can be answered by a machine on its own. That is the whole reason this is a context rather than a set of node-local modules — and it is exactly what the current arrangement gets wrong, by computing all five on the node from a direct database connection.

The shape: decided centrally, delivered as files

Every one of the five resolves the same way, and it is worth stating once rather than five times:

The connectivity context computes the configuration. It arrives over the link as file resources. The service reads files and knows nothing about the mesh.

This costs no new host vocabulary. file, directory, service and container already exist; WireGuard, the resolver and the proxy are all a container or a package, plus files.

It is also what removes the last two upward dependencies. Research 006 counted exactly two modules opening a direct connection to the control plane's database — wireguard and traefik — and they are the reason every node permanently holds a credential to it (ADR 0039). Both are connectivity modules. Closing this context closes that set.

The order it comes up in

The one thing to get right, because everything else depends on it:

0  the node has an underlay address     the machine's own — DHCP, or a provider gave it one
1  the node dials the mesh              OVER THE UNDERLAY, at the address in its token
2  it proves itself, and is proved to   the link exists      (ADR 0039, ADR 0051)
3  the mesh grants it an identity       and an overlay address
4  the overlay comes up                 peer graph delivered as files
5  names resolve                        resolver config delivered as files
6  filtering is applied                 derived from what is assigned here
7  routes and certificates              once this node has something to expose

Step 1 runs on the underlay and never on the overlay. This is the circularity that must not be created: the overlay is configured by the mesh, so a link that required the overlay could never be established on a new node. The link stays on the underlay permanently — it is outbound-only and carries its own identity, so it needs nothing the overlay provides.

Nothing before step 3 can resolve a mesh name, which is why the token carries an address (ADR 0051). Today this is patched with an /etc/hosts floor written underneath the resolver; under this design there is nothing to patch.

1 — The overlay

What is decided: the peer graph. For every node: its overlay address, which peers it holds, which of those it may dial, and which must dial it.

Inputs, all declared:

  • reachability — an endpoint, or none (ADR 0050). Not inferred from the address shape, which is wrong for carrier-grade NAT, wrong for IPv6, and wrong for a routable address behind a closed firewall.
  • site — where the machine physically is, or nothing if it roams.
  • role — hub or not, declared. Today it is inferred from an address prefix, which means a renumbering is an outage and nothing can be asked which node is the hub.

Keys. Each node generates its own keypair. The private key never leaves the machine; the public key is published to the mesh. This is already true and it is already right — it is ADR 0039's a node holds its own identity applied to the overlay, and it means the control plane computes a graph it cannot itself impersonate.

Shape: a hub, with direct peering between co-located nodes.

two nodes at the same site peer directly, host-routed, with a keepalive
everything else routes through the hub
a node with no site — it roams hub only

Roaming is hub-only deliberately, and the reason is a property of WireGuard rather than a preference: there is no failover. A more specific route to a dead endpoint blackholes; it does not fall back to the general one. So a node whose location changes gets exactly one path, because two paths would mean one of them silently swallowing traffic.

What the host receives: an interface configuration and a peer list, as files. It does not compute them, and after this it holds no credential to the mesh's database.

2 — Resolution

Two name spaces, and they do not mix:

resolves to certified by
internal names overlay addresses the mesh CA
public names whatever the outside world must reach a public authority

A node's mesh name is its overlay address. Its public name, if it has one, is a separate fact used by things outside the mesh — and the separation carries two lessons that were learned expensively enough to be worth restating:

  • Mesh names are not multicast names. A name resolved by local multicast discovery introduces a delay and a failure mode that appears on one node and not others — the worst shape a fault can have.
  • A node must not pin its own public name locally. The duplicate record breaks resolution of that name for everything else that needs it.

What the host receives: the resolver's configuration, as files, listing every peer's internal name and overlay address.

What goes away: the /etc/hosts floor. It exists because a node had to reach the mesh database before its own DNS existed; with ADR 0051 nothing needs a name before the link, and a fallback nothing needs is a path nothing tests.

3 — Exposure

Settled by ADR 0049; summarised here because this is where it belongs.

A route is a grant. A module that must be reachable declares it needs one; the proxy provides it and hands back the public name. Ordinary ADR 0044 vocabulary — the mirror of a database grant, where the consumer supplies a target and receives a name rather than supplying nothing and receiving credentials.

A workload on an unreachable node is proxied by a reachable one, across the overlay. Which is the case is a mesh-level fact, which is the fourth reason exposure is control-plane work.

4 — Filtering

Derived from what is assigned here, and from the overlay's shape — a node's open ports are a consequence of what runs on it and who must reach it, not an independent declaration to keep in step by hand.

A rule names its source (ADR 0052). A rule with no source is open, and must say so rather than appear to restrict something. scope: is removed rather than implemented: five manifests carry it today, it is referenced by no code, and it is the clearest instance in the repository of an unenforced rule is indistinguishable from a wrong one, and costs more, because people believe it.

Unknown keys are refused — the discipline the host's declaration parser already has (ADR 0043), and the one manifests lack. scope: survived because nothing rejected it.

5 — Certificates

Two authorities, kept separate on purpose.

issued by for
public names a public ACME authority anything outside the mesh reaches
internal names the mesh CA node-to-node, over the overlay

The split is not collapsed, including in the lab. A single-CA lab would hide any bug living in the split, so the lab runs its own ACME issuer on its public segment and keeps the mesh CA unchanged (research 004).

Public issuance requires genuine public reachability. The HTTP-01 challenge must be answered at the name being certified, so issuance happens through a publicly reachable node regardless of where the workload runs — the same asymmetry as exposure, for the same reason.

The issuer must be configurable. Today it is not: the proxy sets no caServer and therefore defaults to the public authority's production endpoint. Two consequences, and the second is worse than the lab problem that found it — every certificate experiment on a real node consumes production issuance quota, and a retry loop can exhaust it for a week.

The mesh CA is not a bootstrap concern. A joining node verifies the control plane against the fingerprint in its token (ADR 0051), so nothing needs the CA before membership. It certifies internal names afterwards, and that is all it does.

What this removes

The list is worth having in one place, because it is most of the argument:

  • The last two direct database connections from nodes — wireguard and traefik, the only two, both connectivity.
  • Therefore the database credential on every node, and the object-store credential beside it. ADR 0039's central claim becomes true rather than aspirational.
  • The /etc/hosts floor, and the bootstrap circularity it patched.
  • Hub election by address prefix, and the silent no-hub failure when nobody knew the convention.
  • The RFC1918 inference, and the lab substitution that existed to satisfy it.
  • scope:, and the class of manifest key that means nothing.

Open

  • What happens when the hub is down. Resolved by ADR 0053, together with 06's matching question — they were one question. Nothing takes over. WireGuard has no failover, the hub is declared rather than elected, and non-co-located paths stop while co-located direct peers and every already-assigned workload keep running. The recovery path is restore, and its deadline is certificate renewal.
  • Renumbering the overlay. Made possible by declaring the hub rather than inferring it from an address, but no procedure exists, and a graph delivered node by node has an ordering problem while it is half-applied.
  • Revoking a route when a module is unassigned (ADR 0049). A stale public name pointing at nothing fails more visibly than a stale grant.
  • IPv6. ADR 0050 makes it expressible; nothing here says the overlay or the resolver handle it.
  • Reporting declared-versus-observed. ADR 0050 makes the disagreement detectable and does not say who looks or what they are told.