Files
hq/02-DECISIONS/0007-connectivity.md
jschoubben 918dc04916 What this actually is, and three things that were assumed
Four things settled by talking them through, all of which had been true in
somebody's head and written nowhere.

It is not a mesh in the peer-to-peer sense and will not become one. 0001 now
says what it is instead: machines linked by a private network, one node holding
knowledge of all of them, modules as the way anything is built and delivered,
and agents hired onto nodes to do the work. The word describes what machines
can reach, not how they are governed. "Master" overstates it the other way --
nothing needs that node to keep running, only to change.

0006 gains the option that would make it a real mesh, recorded as considered
rather than rejected by silence: every node holding the whole inventory, a
replication process, an elected master with promotion on failure. What settles
it is not the complexity but that it still would not deliver the name, because
application databases are not replicated -- so a genuine peer-to-peer mesh
means becoming a replicated database system for every consumer's data too. That
is a larger product than the thing it would support.

Also in 0006: three central roles, not one. Losing the control plane costs
change, losing the broker costs being told anything, and losing the hub costs
nodes in different places reaching each other at all -- which is operation, not
administration. Whether they are one node is not decided.

And SSH access is identity's. It appeared three times as something that uses
the overlay and never as something the mesh provides, which reads as settled
when nothing decided it. Nobody else could: the mesh is the only thing that
knows which humans and agents exist and which nodes they may reach. Node to
node SSH stays out -- the host has no inbound control surface by decision, and
nodes reaching each other that way is a second control path through the back
door.

0007 gains the requirement underneath all of it. Reachability was recorded as a
fact to track and never as a thing some node must have. The broker's node and
the hub must be dialable by every node at a stable address, or nothing can join
and a disconnected node cannot return. A mesh entirely behind NAT cannot be
raised. That is a precondition and it belongs with the others.

The link staying on the underlay is also argued now rather than asserted. At
join time it is forced; afterwards it is a choice, and the reason is that a
repair channel carried over the thing being repaired is not one. Moving it onto
the overlay, with fallback, is recorded as open with what it would have to get
right -- a WireGuard interface has no link state to test, and a silent fallback
is this repository's recurring fault in a new place.

0010 says in one line what was the intention throughout: the module system is
the CI/CD. Not a pipeline beside the mesh. Build, test, publish and deploy are
one reconciliation seen at four points, which is why a thing that cannot be a
module cannot be delivered.
2026-08-29 13:01:36 +02:00

192 lines
9.9 KiB
Markdown

---
topic: the tiers
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 7. Connectivity
*Consolidated 2026-08-28 from three records. Overlay, resolution, exposure, filtering and
certificates are one design.*
## Why it is control-plane work
Apply the test — *everything that needs to know about more than one node* — and not one of the
five can be answered by a machine on its own:
| | needs to know |
|---|---|
| **overlay** — who peers with whom | every node, and which can be dialled |
| **resolution** — which name is which node | every node |
| **exposure** — which public name reaches which container | which node is publicly reachable |
| **filtering** — which port is open, to whom | what is assigned here, and the overlay's shape |
| **certificates** — who may present which name | which name belongs to which node |
That is exactly what the current arrangement gets wrong, by computing all five on the node from a
direct database connection. Two modules do this, and they are the only two left holding a
credential to the control plane's database.
**The shape of the fix, once for all five:** the connectivity context computes the configuration;
it arrives over the link as `file` resources; the service reads files and knows nothing about the
mesh. **This costs no new host vocabulary.**
## A route is a grant
**Ingress is not substrate.** The control plane does not need a route to start — it listens
locally — and no node needs one to reach it, because the node dials out and has no listening
control surface. It grants itself a route afterwards, the way it grants itself a bucket.
The strongest objection deserves stating: the `api` is the one interface every surface speaks to,
so eventually it *does* want a public name. But **wanting one later is not needing one to
start**, and that distinction is the entire substrate test.
**A module that must be reachable declares it needs a route; the proxy provides one.** Ordinary
instantiation, with the direction mirrored — the consumer supplies a target and receives a name.
**Exposure is three facts at two scopes**, which is why it cannot live on the node:
| the fact | scope |
|---|---|
| the public name resolves to an address | **mesh** — which node is publicly reachable |
| a certificate valid for that name exists | **mesh** — issued once, used on one node |
| the proxy maps that name to that container | **node** |
**A node without a public address is proxied by one that has**, across the overlay. Most nodes sit
behind a connection with no forwarded port, so exposure cannot assume the workload's node is
reachable.
## Reachability is declared, not inferred
The overlay's peer graph is computed from whether a node can be dialled, and that was inferred
from a regular expression over the address. **The address is evidence of reachability; it is not
the fact**, and the gap has already cost:
| address | the regex says | actually |
|---|---|---|
| `100.64.0.0/10` — carrier-grade NAT | **public** | **not reachable.** An endpoint is written to an address nothing can reach |
| any IPv6 address | public | the test is v4 shapes only |
| a routable address behind a closed firewall | public | not reachable |
| a documentation range standing in for a public segment | private | reachable — this is the lab bug |
**A test environment having to choose its addresses to satisfy a regex is the regex telling us it
is not a fact.**
So: **an endpoint, or none** — declared. And **the hub is declared, never derived from an address
prefix**, because an election decided by the first four characters of an address fails silently,
cannot be queried, and makes a renumbering an outage.
**The address remains evidence and stops being the fact.** Where an observed endpoint disagrees
with a declared one, the disagreement is a **reportable condition**, not a silent correction.
**What does not change** is the lesson underneath: role does not imply reachability — a
home-hosted node is a server that cannot be dialled. This keeps that and stops encoding it as a
pattern match.
### Some nodes must be reachable, and this had not been said
*Written 2026-08-29, on being asked and finding no answer.*
Everything above treats reachability as a **fact to record** — which node can be dialled, so that
exposure and certificate issuance can be placed. It never said the converse, and the converse is a
hard requirement:
| role | dialled by | so it needs |
|---|---|---|
| **the node running the broker** | every node, outbound ([ADR 0002](0002-nodes-communicate-over-a-broker.md)) | to be reachable from wherever nodes are, at a **stable address** |
| **the hub** | every node not co-located with its peer | the same |
**Across the internet, "reachable from wherever nodes are" means publicly reachable.** For a mesh
confined to one network it does not — the requirement is about the nodes that exist, not about the
public internet.
**Stable is the sharper half.** A token carries the broker's *address, not a name*, because there
is no resolution before joining ([ADR 0004](0004-a-node-and-how-it-joins.md)). A broker node whose
address moves invalidates every token issued for it, and a node that was disconnected across the
change cannot get back.
**A mesh whose nodes are all behind NAT cannot be raised.** That is a real precondition and it
belongs with the others rather than being discovered.
### The link stays on the underlay, and that is a repair channel
The obvious objection is that all traffic should run over the overlay. Nearly all of it does — SSH,
services, node to node — and the exception is each node's own outbound link to the broker.
**At join time it is forced**: a node has no overlay yet, so it cannot use one to ask for one.
**Afterwards it is a choice**, and the reason is that the link is how a broken node is fixed. A
repair channel carried over the thing being repaired is not a repair channel: a node whose only
path home was the overlay is gone the moment an overlay declaration is wrong.
**What it does not cost is confidentiality.** The link is already authenticated and encrypted
against a pinned fingerprint ([ADR 0004](0004-a-node-and-how-it-joins.md)), so moving it onto the
overlay would not protect traffic that is unprotected today.
## A filter rule names its source
`scope: public` is declared in five manifests, is part of no rule type, and is **referenced by no
code**. So five manifests appear to restrict a port and restrict nothing — on the modules most
worth restricting.
**A rule names its source. `from:` is the only way to scope one, and a rule without one is open**
— which it must say plainly rather than appear to deny.
**`scope:` is removed rather than implemented**, because giving it meaning would leave two ways to
express one thing. And the general fix is that **an unknown key is refused**: the host's
declaration parser already works this way, and manifests are the layer where that discipline is
missing. `scope:` survived because nothing rejected it, and it spread by copying to five
manifests.
## Order, and what it costs
**The link runs on the underlay and never on the overlay.** The overlay is configured by the mesh,
so a link requiring it could never be established on a new node.
**The first declaration is the overlay and nothing else** — because a node's address and peers are
*assigned* so it cannot come earlier, and because it is the way back in. A node reachable over the
overlay can be fixed by hand if a later declaration breaks the machine; **a large first
declaration risks a node that is broken and unreachable at once.**
**Reachable is not the same as having a control surface.** Every node reaches every other over the
overlay — SSH, services, ordinary traffic — and every node consumes from the broker. What is
forbidden is a listening thing that accepts instructions and changes the machine.
## Consequences
- **The last two direct database connections leave the nodes**, and with them the database
credential every node carries.
- **The `/etc/hosts` floor goes**, along with the bootstrap circularity it patched.
- **Two certificate authorities stay separate on purpose**: a public one for public names, the
mesh's own for internal ones. A single-CA lab would hide any bug living in the split.
- **What happens when the hub is down**: nothing takes over. Non-co-located paths stop; co-located
peers and every assigned workload keep running.
## Open — the link over the overlay, with a fallback
*Raised 2026-08-29 and deliberately left open, because the honest gain is smaller than it looks
and it is a decision rather than a derivation.*
The proposal: a node prefers the overlay for its link and drops to the underlay when the overlay
is not working — so ordinary operation is private and the underlay stays as the way back.
**Two things it would have to get right:**
- **The trigger cannot be "is the overlay up".** A WireGuard interface has no link state; once
configured it is up whether or not the far end exists. So there is no flag to read, and failing
over means *try, fail, time out, retry elsewhere*.
- **Running on the fallback has to be visible.** A node that quietly drops to the underlay is a
node whose overlay is broken with nothing to say so, and it will stay broken because everything
still works. That is this repository's recurring fault — a failure that reads as success — and a
fallback is the easiest place in the design to reintroduce it.
**What stops it being an obvious win:** if the underlay path must stay available for the fallback,
the broker stays exposed on it. So the exposure is unchanged and the traffic was already encrypted
— the gain is which network carries bytes, not what an attacker can reach.
**The version that would buy something is overlay-only**, with the broker firewalled to the overlay
in steady state, accepting that a node whose overlay breaks needs hands-on recovery. That is a real
trade: it exchanges the automatic way back in for a closed port.
Not decided either way here.