--- topic: the tiers status: accepted date: 2026-08-28 deciders: jochen reconstructed: false --- # 7. Connectivity *Consolidated 2026-08-28 from three records. Overlay, resolution, exposure, filtering and certificates are one design.* ## Why it is control-plane work Apply the test — *everything that needs to know about more than one node* — and not one of the five can be answered by a machine on its own: | | needs to know | |---|---| | **overlay** — who peers with whom | every node, and which can be dialled | | **resolution** — which name is which node | every node | | **exposure** — which public name reaches which container | which node is publicly reachable | | **filtering** — which port is open, to whom | what is assigned here, and the overlay's shape | | **certificates** — who may present which name | which name belongs to which node | That is exactly what the current arrangement gets wrong, by computing all five on the node from a direct database connection. Two modules do this, and they are the only two left holding a credential to the control plane's database. **The shape of the fix, once for all five:** the connectivity context computes the configuration; it arrives over the link as `file` resources; the service reads files and knows nothing about the mesh. **This costs no new host vocabulary.** ## A route is a grant **Ingress is not substrate.** The control plane does not need a route to start — it listens locally — and no node needs one to reach it, because the node dials out and has no listening control surface. It grants itself a route afterwards, the way it grants itself a bucket. The strongest objection deserves stating: the `api` is the one interface every surface speaks to, so eventually it *does* want a public name. But **wanting one later is not needing one to start**, and that distinction is the entire substrate test. **A module that must be reachable declares it needs a route; the proxy provides one.** Ordinary instantiation, with the direction mirrored — the consumer supplies a target and receives a name. **Exposure is three facts at two scopes**, which is why it cannot live on the node: | the fact | scope | |---|---| | the public name resolves to an address | **mesh** — which node is publicly reachable | | a certificate valid for that name exists | **mesh** — issued once, used on one node | | the proxy maps that name to that container | **node** | **A node without a public address is proxied by one that has**, across the overlay. Most nodes sit behind a connection with no forwarded port, so exposure cannot assume the workload's node is reachable. ## Reachability is declared, not inferred The overlay's peer graph is computed from whether a node can be dialled, and that was inferred from a regular expression over the address. **The address is evidence of reachability; it is not the fact**, and the gap has already cost: | address | the regex says | actually | |---|---|---| | `100.64.0.0/10` — carrier-grade NAT | **public** | **not reachable.** An endpoint is written to an address nothing can reach | | any IPv6 address | public | the test is v4 shapes only | | a routable address behind a closed firewall | public | not reachable | | a documentation range standing in for a public segment | private | reachable — this is the lab bug | **A test environment having to choose its addresses to satisfy a regex is the regex telling us it is not a fact.** So: **an endpoint, or none** — declared. And **the hub is declared, never derived from an address prefix**, because an election decided by the first four characters of an address fails silently, cannot be queried, and makes a renumbering an outage. **The address remains evidence and stops being the fact.** Where an observed endpoint disagrees with a declared one, the disagreement is a **reportable condition**, not a silent correction. **What does not change** is the lesson underneath: role does not imply reachability — a home-hosted node is a server that cannot be dialled. This keeps that and stops encoding it as a pattern match. ### Some nodes must be reachable, and this had not been said *Written 2026-08-29, on being asked and finding no answer.* Everything above treats reachability as a **fact to record** — which node can be dialled, so that exposure and certificate issuance can be placed. It never said the converse, and the converse is a hard requirement: | role | dialled by | so it needs | |---|---|---| | **the node running the broker** | every node, outbound ([ADR 0002](0002-nodes-communicate-over-a-broker.md)) | to be reachable from wherever nodes are, at a **stable address** | | **the hub** | every node not co-located with its peer | the same | **Across the internet, "reachable from wherever nodes are" means publicly reachable.** For a mesh confined to one network it does not — the requirement is about the nodes that exist, not about the public internet. **Stable is the sharper half.** A token carries the broker's *address, not a name*, because there is no resolution before joining ([ADR 0004](0004-a-node-and-how-it-joins.md)). A broker node whose address moves invalidates every token issued for it, and a node that was disconnected across the change cannot get back. **A mesh whose nodes are all behind NAT cannot be raised.** That is a real precondition and it belongs with the others rather than being discovered. ### The link stays on the underlay, and that is a repair channel The obvious objection is that all traffic should run over the overlay. Nearly all of it does — SSH, services, node to node — and the exception is each node's own outbound link to the broker. **At join time it is forced**: a node has no overlay yet, so it cannot use one to ask for one. **Afterwards it is a choice**, and the reason is that the link is how a broken node is fixed. A repair channel carried over the thing being repaired is not a repair channel: a node whose only path home was the overlay is gone the moment an overlay declaration is wrong. **What it does not cost is confidentiality.** The link is already authenticated and encrypted against a pinned fingerprint ([ADR 0004](0004-a-node-and-how-it-joins.md)), so moving it onto the overlay would not protect traffic that is unprotected today. ## A filter rule names its source `scope: public` is declared in five manifests, is part of no rule type, and is **referenced by no code**. So five manifests appear to restrict a port and restrict nothing — on the modules most worth restricting. **A rule names its source. `from:` is the only way to scope one, and a rule without one is open** — which it must say plainly rather than appear to deny. **`scope:` is removed rather than implemented**, because giving it meaning would leave two ways to express one thing. And the general fix is that **an unknown key is refused**: the host's declaration parser already works this way, and manifests are the layer where that discipline is missing. `scope:` survived because nothing rejected it, and it spread by copying to five manifests. ## Order, and what it costs **The link runs on the underlay and never on the overlay.** The overlay is configured by the mesh, so a link requiring it could never be established on a new node. **The first declaration is the overlay and nothing else** — because a node's address and peers are *assigned* so it cannot come earlier, and because it is the way back in. A node reachable over the overlay can be fixed by hand if a later declaration breaks the machine; **a large first declaration risks a node that is broken and unreachable at once.** **Reachable is not the same as having a control surface.** Every node reaches every other over the overlay — SSH, services, ordinary traffic — and every node consumes from the broker. What is forbidden is a listening thing that accepts instructions and changes the machine. ## Consequences - **The last two direct database connections leave the nodes**, and with them the database credential every node carries. - **The `/etc/hosts` floor goes**, along with the bootstrap circularity it patched. - **Two certificate authorities stay separate on purpose**: a public one for public names, the mesh's own for internal ones. A single-CA lab would hide any bug living in the split. - **What happens when the hub is down**: nothing takes over. Non-co-located paths stop; co-located peers and every assigned workload keep running. ## Open — the link over the overlay, with a fallback *Raised 2026-08-29 and deliberately left open, because the honest gain is smaller than it looks and it is a decision rather than a derivation.* The proposal: a node prefers the overlay for its link and drops to the underlay when the overlay is not working — so ordinary operation is private and the underlay stays as the way back. **Two things it would have to get right:** - **The trigger cannot be "is the overlay up".** A WireGuard interface has no link state; once configured it is up whether or not the far end exists. So there is no flag to read, and failing over means *try, fail, time out, retry elsewhere*. - **Running on the fallback has to be visible.** A node that quietly drops to the underlay is a node whose overlay is broken with nothing to say so, and it will stay broken because everything still works. That is this repository's recurring fault — a failure that reads as success — and a fallback is the easiest place in the design to reintroduce it. **What stops it being an obvious win:** if the underlay path must stay available for the fallback, the broker stays exposed on it. So the exposure is unchanged and the traffic was already encrypted — the gain is which network carries bytes, not what an attacker can reach. **The version that would buy something is overlay-only**, with the broker firewalled to the overlay in steady state, accepting that a node whose overlay breaks needs hands-on recovery. That is a real trade: it exchanges the automatic way back in for a closed port. Not decided either way here.