Files
hq/02-DECISIONS/0007-connectivity.md
T
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00

6.1 KiB

status, date, deciders, reconstructed
status date deciders reconstructed
accepted 2026-08-28 jochen false

7. Connectivity

Consolidated 2026-08-28 from three records. Overlay, resolution, exposure, filtering and certificates are one design.

Why it is control-plane work

Apply the test — everything that needs to know about more than one node — and not one of the five can be answered by a machine on its own:

needs to know
overlay — who peers with whom every node, and which can be dialled
resolution — which name is which node every node
exposure — which public name reaches which container which node is publicly reachable
filtering — which port is open, to whom what is assigned here, and the overlay's shape
certificates — who may present which name which name belongs to which node

That is exactly what the current arrangement gets wrong, by computing all five on the node from a direct database connection. Two modules do this, and they are the only two left holding a credential to the control plane's database.

The shape of the fix, once for all five: the connectivity context computes the configuration; it arrives over the link as file resources; the service reads files and knows nothing about the mesh. This costs no new host vocabulary.

A route is a grant

Ingress is not substrate. The control plane does not need a route to start — it listens locally — and no node needs one to reach it, because the node dials out and has no listening control surface. It grants itself a route afterwards, the way it grants itself a bucket.

The strongest objection deserves stating: the api is the one interface every surface speaks to, so eventually it does want a public name. But wanting one later is not needing one to start, and that distinction is the entire substrate test.

A module that must be reachable declares it needs a route; the proxy provides one. Ordinary instantiation, with the direction mirrored — the consumer supplies a target and receives a name.

Exposure is three facts at two scopes, which is why it cannot live on the node:

the fact scope
the public name resolves to an address mesh — which node is publicly reachable
a certificate valid for that name exists mesh — issued once, used on one node
the proxy maps that name to that container node

A node without a public address is proxied by one that has, across the overlay. Most nodes sit behind a connection with no forwarded port, so exposure cannot assume the workload's node is reachable.

Reachability is declared, not inferred

The overlay's peer graph is computed from whether a node can be dialled, and that was inferred from a regular expression over the address. The address is evidence of reachability; it is not the fact, and the gap has already cost:

address the regex says actually
100.64.0.0/10 — carrier-grade NAT public not reachable. An endpoint is written to an address nothing can reach
any IPv6 address public the test is v4 shapes only
a routable address behind a closed firewall public not reachable
a documentation range standing in for a public segment private reachable — this is the lab bug

A test environment having to choose its addresses to satisfy a regex is the regex telling us it is not a fact.

So: an endpoint, or none — declared. And the hub is declared, never derived from an address prefix, because an election decided by the first four characters of an address fails silently, cannot be queried, and makes a renumbering an outage.

The address remains evidence and stops being the fact. Where an observed endpoint disagrees with a declared one, the disagreement is a reportable condition, not a silent correction.

What does not change is the lesson underneath: role does not imply reachability — a home-hosted node is a server that cannot be dialled. This keeps that and stops encoding it as a pattern match.

A filter rule names its source

scope: public is declared in five manifests, is part of no rule type, and is referenced by no code. So five manifests appear to restrict a port and restrict nothing — on the modules most worth restricting.

A rule names its source. from: is the only way to scope one, and a rule without one is open — which it must say plainly rather than appear to deny.

scope: is removed rather than implemented, because giving it meaning would leave two ways to express one thing. And the general fix is that an unknown key is refused: the host's declaration parser already works this way, and manifests are the layer where that discipline is missing. scope: survived because nothing rejected it, and it spread by copying to five manifests.

Order, and what it costs

The link runs on the underlay and never on the overlay. The overlay is configured by the mesh, so a link requiring it could never be established on a new node.

The first declaration is the overlay and nothing else — because a node's address and peers are assigned so it cannot come earlier, and because it is the way back in. A node reachable over the overlay can be fixed by hand if a later declaration breaks the machine; a large first declaration risks a node that is broken and unreachable at once.

Reachable is not the same as having a control surface. Every node reaches every other over the overlay — SSH, services, ordinary traffic — and every node consumes from the broker. What is forbidden is a listening thing that accepts instructions and changes the machine.

Consequences

  • The last two direct database connections leave the nodes, and with them the database credential every node carries.
  • The /etc/hosts floor goes, along with the bootstrap circularity it patched.
  • Two certificate authorities stay separate on purpose: a public one for public names, the mesh's own for internal ones. A single-CA lab would hide any bug living in the split.
  • What happens when the hub is down: nothing takes over. Non-co-located paths stop; co-located peers and every assigned workload keep running.