Settles the design repository now that the self-upgrade build is on main: - Records the two decisions that shipped without a record — ADR 0077 (the controller/foundation/node vocabulary) and ADR 0078 (the store and broker are ordinary modules); accepts ADR 0075 and 0076, which shipped work rests on. - Fills issue 051's amended-design and wires ADR 0078 into 07-the-foundation. - Sweeps the repo rename (mesh-control -> mesh-controller) into the mutable docs now that the forge repo is renamed; updates the glossary note and repos.md. - Fixes the six broken links from the design-doc renames, indexes the glossary, regenerates the decisions reading order. Both checks (records.py, index.py) are green. Statuses stay honest: the build is on main and lab-proven but not deployed as the production mesh, so the to-be docs remain in-progress and the as-is layer (the hal mesh) is unchanged — graduation to implemented + as-is belongs to deployment, not merge. https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
645 lines
39 KiB
Markdown
645 lines
39 KiB
Markdown
---
|
|
layer: to-be
|
|
status: in-progress
|
|
code:
|
|
- mesh-controller internal/catalogue/filtering.go
|
|
- mesh-controller examples/route-proxy
|
|
- mesh-controller internal/identity/authority.go
|
|
- mesh-host internal/identity/serving.go
|
|
- mesh-host internal/apply (the service that reflects a rule set)
|
|
updated: 2026-09-09
|
|
decisions:
|
|
- 02-DECISIONS/0005-the-node-host.md
|
|
- 02-DECISIONS/0004-a-node-and-how-it-joins.md
|
|
- 02-DECISIONS/0007-connectivity.md
|
|
- 02-DECISIONS/0007-connectivity.md
|
|
- 02-DECISIONS/0004-a-node-and-how-it-joins.md
|
|
- 02-DECISIONS/0007-connectivity.md
|
|
- 02-DECISIONS/0006-the-substrate-and-the-control-plane.md
|
|
- 02-DECISIONS/0066-public-routing-is-name-agnostic.md
|
|
---
|
|
|
|
# Connectivity
|
|
|
|
One of [the controller's](06-the-controller.md) ten contexts, and the one with the most
|
|
moving parts: **overlay, resolution, exposure, filtering, certificates.**
|
|
|
|
It is written as a whole because the five are one design. They share inputs, they must agree, and
|
|
every one of them today is computed in a different place by a different module from a different
|
|
copy of the same facts.
|
|
|
|
## Why it is controller work
|
|
|
|
Apply [the test](06-the-controller.md) — *everything that needs to know about more than one
|
|
node* — to each responsibility:
|
|
|
|
| | needs to know | whose |
|
|
|---|---|---|
|
|
| **overlay** — who peers with whom, at what address | **every node**, and which of them can be dialled | controller |
|
|
| **resolution** — which name is which node | **every node** | controller |
|
|
| **exposure** — which public name reaches which container | **which node is publicly reachable** ([ADR 0007](../../02-DECISIONS/0007-connectivity.md)) | controller |
|
|
| **filtering** — which port is open, to whom | what is assigned here, and the overlay's shape | controller decides, host applies |
|
|
| **certificates** — who may present which name | which name belongs to which node | controller |
|
|
|
|
**Not one of the five can be answered by a machine on its own.** That is the whole reason this is
|
|
a context rather than a set of node-local modules — and it is exactly what the current
|
|
arrangement gets wrong, by computing all five on the node from a direct database connection.
|
|
|
|
## The shape: decided centrally, delivered as files
|
|
|
|
Every one of the five resolves the same way, and it is worth stating once rather than five times:
|
|
|
|
> **The connectivity context computes the configuration. It arrives over the link as `file`
|
|
> resources. The service reads files and knows nothing about the mesh.**
|
|
|
|
This costs **no new host vocabulary**. `file`, `directory`, `service` and `container` already
|
|
exist; WireGuard, the resolver and the proxy are all *a container or a package, plus files*.
|
|
|
|
It is also what removes the last two upward dependencies.
|
|
[Research 006](../../01-RESEARCH/006-mesh-from-scratch/host-size.md) counted exactly two modules
|
|
opening a direct connection to the controller's database — `wireguard` and `traefik` — and
|
|
they are the reason every node permanently holds a credential to it
|
|
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). Both are connectivity
|
|
modules. **Closing this context closes that set.**
|
|
|
|
### And they are modules, not a second mechanism beside the module system
|
|
|
|
*Written 2026-08-29, from building it. The first version was code beside the module system doing
|
|
the module system's job, and the fault it produced is the point of writing this down.*
|
|
|
|
**A machine was on the private network because it had an address.** Every node that had been
|
|
placed got a peer list, whether or not anybody wanted it there, and there was no way to say a
|
|
machine should stay off. That is what "special-cased" cost, and it was invisible until somebody
|
|
wanted the exception.
|
|
|
|
**What made it look unavoidable:** a peer list cannot be written in a manifest. It is derived from
|
|
every other machine, so it differs on each one and changes when any of them changes. So the
|
|
manifest says its resources are **computed** — it names something in the controller that works
|
|
them out per node — and it is a module in every other respect: assigned, resolved, configured by
|
|
settings, and absent from a machine nobody gave it to.
|
|
|
|
**What that made possible immediately** is the arrangement below, which the code has:
|
|
|
|
| module | provides | requires | claims |
|
|
|---|---|---|---|
|
|
| the WireGuard one | a private network, **and the mesh's own addressing** | | *the* private network, one per node |
|
|
| the names one | name resolution | the mesh's own addressing | |
|
|
| `networking` | | both of the above | |
|
|
|
|
**Three rather than one, because WireGuard is one VPN of several.** Naming the module after the
|
|
job — `networking` — and putting WireGuard inside it is the retired *flavor* idea wearing a
|
|
generic name: the second VPN has nowhere to go. So a module is named for what it *is* and declares
|
|
what it *does*, and `networking` is the third row — requirements and no files
|
|
([ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md)).
|
|
|
|
**Names left the WireGuard module for their own.** They had been delivered inside it, on the
|
|
argument that a machine with peers and no names is half on the network. True, and the wrong place
|
|
to fix it — names are identical over a *different* private network, so bundling them made one
|
|
module out of two things. They require the mesh's **addressing** rather than a private network in
|
|
general, because that is what they are computed from: over a VPN that hands out its own addresses
|
|
the mesh has nothing to write, and refusing is what stops a machine being given a hosts file that
|
|
means nothing on it.
|
|
|
|
**And the claim is not decoration.** Choosing a different VPN still installed WireGuard — dragged
|
|
back in by the names, which needed addresses only WireGuard hands out — and nobody was told.
|
|
Running two VPNs is not always wrong; being *the* one the mesh runs over is singular. So it is a
|
|
claim, and the collision is refused by name.
|
|
|
|
**And the proxy's half, which was the other module reaching into the database.** A web application
|
|
requiring a reverse proxy has to say *which name, which port*, and there was nowhere to put it —
|
|
`requires` says a thing must exist and never said what to do with it. A module now contributes to
|
|
a requirement, the controller collects every contribution on a node, and the provider is given
|
|
them as a file at a path it named. It reloads when that file changes, by the same `restart-on` the
|
|
private network needed when a peer list changed under a running interface.
|
|
|
|
**The proxy's configuration is not written by the mesh.** It is given the facts and turns them
|
|
into whatever it runs, which is why swapping Traefik for something else touches nothing that
|
|
publishes through it. See [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) for the
|
|
other direction — handing a credential *back* — which is the larger half and is not built.
|
|
|
|
**What is still not a module, and why that is correct.** The host needs none of this. It has an
|
|
address and a route before the mesh exists — that is the machine's own networking — and the
|
|
broker's address is carried in the token rather than resolved
|
|
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). **The one connection that
|
|
carries modules cannot itself be one.** Everything above it can be, and now is.
|
|
|
|
## The order it comes up in
|
|
|
|
The one thing to get right, because everything else depends on it:
|
|
|
|
```
|
|
0 the node has an underlay address the machine's own — DHCP, or a provider gave it one
|
|
1 the node dials the mesh OVER THE UNDERLAY, at the address in its token
|
|
2 it proves itself, and is proved to the link exists (ADR 0004, ADR 0004)
|
|
3 the mesh grants it an identity and an overlay address
|
|
4 the overlay comes up peer graph delivered as files
|
|
5 names resolve resolver config delivered as files
|
|
6 filtering is applied derived from what is assigned here
|
|
7 routes and certificates once this node has something to expose
|
|
```
|
|
|
|
**Step 1 runs on the underlay and never on the overlay.** This is the circularity that must not
|
|
be created: the overlay is configured by the mesh, so a link that required the overlay could
|
|
never be established on a new node. The link stays on the underlay permanently — it is
|
|
outbound-only and carries its own identity, so it needs nothing the overlay provides.
|
|
|
|
**Nothing before step 3 can resolve a mesh name**, which is why the token carries an *address*
|
|
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). Today this is
|
|
patched with an `/etc/hosts` floor written underneath the resolver; under this design there is
|
|
nothing to patch.
|
|
|
|
**Step 1 has a precondition this document treated as a fact to record rather than a requirement:
|
|
the broker's node must be dialable by every node, at a stable address, and so must the hub**
|
|
([ADR 0007](../../02-DECISIONS/0007-connectivity.md)). Across the internet that means publicly
|
|
reachable; on one network it does not. A mesh whose nodes are all behind NAT cannot be raised, and
|
|
a broker node whose address moves invalidates every token issued for it.
|
|
|
|
**Whether the link should later move onto the overlay, with the underlay as fallback, is
|
|
[open](../../02-DECISIONS/0007-connectivity.md).** It is a decision rather than a derivation: the
|
|
gain is which network carries bytes, not what an attacker can reach, since the link is already
|
|
encrypted against a pinned fingerprint.
|
|
|
|
## 1 — The overlay
|
|
|
|
**What is decided:** the peer graph. For every node: its overlay address, which peers it holds,
|
|
which of those it may dial, and which must dial it.
|
|
|
|
**Inputs, all declared:**
|
|
|
|
- **reachability** — an endpoint, or none
|
|
([ADR 0007](../../02-DECISIONS/0007-connectivity.md)). Not
|
|
inferred from the address shape, which is wrong for carrier-grade NAT, wrong for IPv6, and
|
|
wrong for a routable address behind a closed firewall.
|
|
- **site** — where the machine physically is, or nothing if it roams.
|
|
- **role** — hub or not, **declared**. Today it is inferred from an address prefix, which means a
|
|
renumbering is an outage and nothing can be asked which node is the hub.
|
|
|
|
**Keys.** Each node generates its own keypair. **The private key never leaves the machine**; the
|
|
public key is published to the mesh. This is already true and it is already right — it is
|
|
[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)'s *a node holds its own
|
|
identity* applied to the overlay, and it means the controller computes a graph it cannot
|
|
itself impersonate.
|
|
|
|
**Shape: a hub, with direct peering between co-located nodes.**
|
|
|
|
| | |
|
|
|---|---|
|
|
| two nodes at the same site | peer **directly**, host-routed, with a keepalive |
|
|
| everything else | routes through the **hub** |
|
|
| a node with no site — it roams | **hub only** |
|
|
|
|
**Roaming is hub-only deliberately, and the reason is a property of WireGuard rather than a
|
|
preference: there is no failover.** A more specific route to a dead endpoint blackholes; it does
|
|
not fall back to the general one. So a node whose location changes gets exactly one path, because
|
|
two paths would mean one of them silently swallowing traffic.
|
|
|
|
**What the host receives:** an interface configuration and a peer list, as files. It does not
|
|
compute them, and after this it holds no credential to the mesh's database.
|
|
|
|
### Four things the lab found, none of them visible from the mesh's own state
|
|
|
|
*Written 2026-08-29, on the first three machines to actually run this.*
|
|
|
|
Each looked like a working network from every angle the mesh can see: the graph was right, the
|
|
files were right, the services were up, and every node reported success.
|
|
|
|
- **A running interface does not re-read its configuration.** A node joins, every existing node's
|
|
peer list changes, each file is replaced — and the service is already running, so nothing
|
|
reloads it. Every existing node keeps a network that no longer exists. The declaration has to
|
|
say the service must *reflect* the file, which is declared state; a command to restart would be
|
|
an action, and the link may not carry one
|
|
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)).
|
|
- **A hub that shares a site with a spoke was emitted twice** — once as a direct peer and once as
|
|
the route of last resort. WireGuard takes one entry per public key, so the interface refuses the
|
|
file. The ordinary shape of a small mesh, and in none of the tests written before it ran.
|
|
- **Two nodes at one site that neither can be dialled must not peer directly.** Nobody opens the
|
|
path, and the direct route is more specific than the hub's, so it wins and blackholes. This
|
|
document's own warning, arriving in its implementation: *a more specific route to a dead
|
|
endpoint blackholes; it does not fall back to the general one.*
|
|
- **The container runtime closes the door the overlay needs.** Docker sets the FORWARD policy to
|
|
DROP, so a hub with `ip_forward` enabled still carries nothing between its spokes. The foundation
|
|
at tier 1 silently breaks the network at tier 2, and nothing in either tier's state says so. The
|
|
hub inserts its own rule above those chains and removes it on the way down.
|
|
|
|
**The pattern in all four:** the mesh's picture of the network was correct and the network did not
|
|
work. That is the argument for the lab in one line — none of these is reachable by reasoning, and
|
|
each was found within minutes of a real machine trying it.
|
|
|
|
## 2 — Resolution
|
|
|
|
**Two name spaces, and they do not mix:**
|
|
|
|
| | resolves to | certified by |
|
|
|---|---|---|
|
|
| **internal names** | overlay addresses | the **mesh CA** |
|
|
| **public names** | whatever the outside world must reach | a **public authority** |
|
|
|
|
A node's mesh name is its overlay address. Its public name, if it has one, is a separate fact
|
|
used by things outside the mesh — and the separation carries two lessons that were learned
|
|
expensively enough to be worth restating:
|
|
|
|
- **Mesh names are not multicast names.** A name resolved by local multicast discovery introduces
|
|
a delay and a failure mode that appears on one node and not others — the worst shape a fault
|
|
can have.
|
|
- **A node must not pin its own public name locally.** The duplicate record breaks resolution of
|
|
that name for everything else that needs it.
|
|
|
|
**What the host receives:** the resolver's configuration, as files, listing every peer's internal
|
|
name and overlay address.
|
|
|
|
**What goes away:** the `/etc/hosts` floor. It exists because a node had to reach the mesh
|
|
database before its own DNS existed; with [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)
|
|
nothing needs a name before the link, and a fallback nothing needs is a path nothing tests.
|
|
|
|
### Names, and what a container can see
|
|
|
|
*2026-08-31, from a container that could not resolve a name every machine could.*
|
|
|
|
Internal names are `<node>.internal` — the suffix is the one IANA reserved in 2024, so a name that
|
|
leaks into a public resolver fails rather than reaching a stranger's machine. They are computed
|
|
centrally, because a name set needs every node at once, and written to each machine's hosts file.
|
|
|
|
**A file rather than a resolver**, and the reasoning holds: it works on every Linux, needs no
|
|
package, and has no failure mode of its own. The stated trigger for a daemon was *names that are
|
|
not one-per-node* — service names, wildcards.
|
|
|
|
**But a container does not inherit the machine's names.** It gets its own hosts file holding only
|
|
its own hostname. So every name the mesh wrote was invisible to the majority of things that need
|
|
one — and *on the machine it always worked*, which is exactly what made it easy to miss. It was
|
|
found by a database client on one node failing to resolve another node, on a mesh where both names
|
|
were correct and present on both machines.
|
|
|
|
**So the mesh gives its names to the containers it declares**, written into each container's own
|
|
hosts file by the runtime. That extends the file decision rather than overturning it. Given by the
|
|
mesh and not chosen by a module: a module that listed the machines would go stale the day one
|
|
joins, and a module that did not would be one whose containers cannot reach anything by name.
|
|
|
|
**The boundary, which is deliberate and worth stating:** *declared* containers. A container
|
|
somebody starts by hand is not the mesh's to configure, and reaching into every container on a
|
|
machine — declared or not — is what a nameserver in `resolv.conf` would be for.
|
|
|
|
### The resolver, built
|
|
|
|
*2026-08-31.* **A service is reached at `<service>.<node>.internal`** — the first label is the
|
|
service, the rest is the node — so what resolves is *anything under a node's name*, going to that
|
|
node. What routes it once it arrives is a proxy's, and stays separate.
|
|
|
|
**The mesh writes the data and runs no daemon.** One wildcard per machine, from the same set that
|
|
writes the hosts file. A resolver is third-party software and runs *on* the mesh rather than being
|
|
*of* it: the mesh has no business shipping one, choosing which one, or knowing its configuration
|
|
language. Swapping dnsmasq for unbound changes that module and nothing in the controller.
|
|
|
|
**Two roles, two claims, because they are different things.** systemd-resolved cannot answer a
|
|
wildcard at all — it routes the mesh's suffix to something that can. Treating serving and asking
|
|
as one role produces a module that cannot work.
|
|
|
|
| | claims | |
|
|
|---|---|---|
|
|
| serving | `the-dns-port` | answers the wildcards |
|
|
| asking | `the-resolver-configuration` | decides what the machine asks |
|
|
|
|
So *which* resolver is not a mesh-wide decision. One machine can use what systemd already owns and
|
|
another can run dnsmasq, and two of either on one machine is refused rather than fought over.
|
|
|
|
**Two things a resolver must not do**, both found by a machine rather than by reasoning:
|
|
|
|
- **Take an address something else holds.** systemd-resolved holds `127.0.0.53` *and* `127.0.0.54`.
|
|
- **Read `resolv.conf` for its upstreams.** Whatever points a machine at the mesh writes the
|
|
resolver's own address there, so it becomes its own upstream and every query it cannot answer
|
|
loops until its receive queue fills. It needs no upstream: only the mesh's suffix is routed to
|
|
it.
|
|
|
|
*Checked on two machines, through the path an application takes — nsswitch, files, then DNS —
|
|
because the module deciding what the machine asks is half of what is being tested and only that
|
|
path goes through it.*
|
|
|
|
**That is now the second reason to want a resolver**, and it is a different one from the trigger
|
|
above:
|
|
|
|
| | |
|
|
|---|---|
|
|
| names that are not one-per-node | a service named under a machine — `postgres.novox.internal` |
|
|
| containers the mesh did not declare | anything a person or another tool starts on a node |
|
|
|
|
### Which resolver is not a question the mesh answers
|
|
|
|
*Written 2026-08-31, after treating it as open when it had been decided two days earlier.*
|
|
|
|
**A resolver takes over `/etc/resolv.conf`, which is a singular resource, so it is a claim** —
|
|
[ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) lists it in the table beside the seat
|
|
and pid 1. Choosing between resolved, dnsmasq and unbound is **assigning a module**, per machine,
|
|
and two of them cannot both be assigned there:
|
|
|
|
> `resolved-config and dnsmasq both claim "/etc/resolv.conf", and only one thing may hold it per node`
|
|
|
|
So there is nothing global to settle and nothing for the mesh to guess. One machine can use what
|
|
systemd already owns and another can run dnsmasq, and neither has to know about the other.
|
|
|
|
**What the mesh contributes is the part only it can know**: which machines exist and where. That
|
|
is `mesh-resolver`, which writes one file and holds no claim, because writing a file takes nothing
|
|
over. A daemon module requires that data and claims the resolver — so swapping the daemon changes
|
|
that module and nothing else.
|
|
|
|
**This was recorded on 2026-08-29 and reopened as an unanswered question on the 31st.** Which is
|
|
the argument for the table in ADR 0009 being a table: the pattern is only obvious once seen, and
|
|
the cost of not seeing it is inventing a mechanism that already exists.
|
|
|
|
### And the public names a proxy serves must resolve in the mesh too
|
|
|
|
*2026-09-09, found by an internal certificate authority that could not issue.* The mesh writes every
|
|
`<node>.internal` name into every declared container and treats the public names a proxy serves as a
|
|
separate matter — *what routes it once it arrives is a proxy's, and stays separate*, above. That
|
|
holds for a client dialling by internal name. It does not hold for anything **inside** the mesh that
|
|
must reach a public name, and the first such thing to appear was the internal issuer of §5.
|
|
|
|
**An issuer validates by connecting to the name it is certifying.** Asked for a certificate for a
|
|
routed public name, the internal authority accepted the order, offered a challenge, and then could
|
|
not connect: nothing in the mesh resolved that name, so the challenge had no target. A name the mesh
|
|
can reach from the outside but cannot resolve from the inside is a name it cannot certify with an
|
|
authority of its own.
|
|
|
|
**So a granted route is published into internal resolution as well** — the routed name to the node
|
|
that serves it, mesh-wide, by the same mechanism that writes the node names. It is *given by the
|
|
mesh, not chosen by a module*, for the same reason the node names are: a module listing the routes
|
|
would go stale the day one changes. The mesh propagates the names it was told to serve and still
|
|
knows nothing about what they mean
|
|
([ADR 0066](../../02-DECISIONS/0066-public-routing-is-name-agnostic.md)).
|
|
|
|
## 3 — Exposure
|
|
|
|
Settled by [ADR 0007](../../02-DECISIONS/0007-connectivity.md); summarised here because
|
|
this is where it belongs.
|
|
|
|
**A route is a grant.** A module that must be reachable declares it needs one; the proxy provides
|
|
it and hands back the public name. Ordinary
|
|
[ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md)
|
|
vocabulary — the mirror of a database grant, where the consumer supplies a target and receives a
|
|
name rather than supplying nothing and receiving credentials.
|
|
|
|
**A workload on an unreachable node is proxied by a reachable one, across the overlay.** Which is
|
|
the case is a mesh-level fact, which is the fourth reason exposure is controller work.
|
|
|
|
### What was built
|
|
|
|
*2026-08-31.* Nothing new in the vocabulary, which was the claim and is now the fact: a route is a
|
|
provision, a proxy provides it, and a module that must be reachable requires it. The consumer
|
|
contributes the name it wants and the port it listens on; the proxy receives every consumer that
|
|
asked; the consumer is told what the provider serves, which is how it knows its own name.
|
|
|
|
**One field was missing, and it is the one anything reaching back needs.** A contribution now
|
|
carries **where the mesh says that machine is**. A database is reached *by* its consumer, so the
|
|
mesh never had to tell a provider where anybody was; a proxy is the other direction — it is told
|
|
to send traffic to a consumer and has to open a connection. Without it every provider implementing
|
|
a provision would have to know how the mesh names machines, which is a convention leaking into
|
|
every module.
|
|
|
|
**Exposure and filtering are different questions and a module answers both.** A workload says what
|
|
it listens on and who may reach it; separately, it says it wants a route. A module that asked for a
|
|
route and not for the port is unreachable by the proxy it just asked for — which the lab
|
|
demonstrates, because the machine is already filtering by the time this runs.
|
|
|
|
**Withdrawal, which was open above.** The file the proxy is given is the whole truth about who has
|
|
a route, so a proxy replaces its table rather than merging. Merging would keep serving a name whose
|
|
module was unassigned — and *a stale public name pointing at nothing fails more visibly than a
|
|
stale grant* is the reason it must not survive, not a reason to tolerate it.
|
|
|
|
**A name a proxy does not serve is refused by saying which it does.** A route that was withdrawn
|
|
and a name that never existed are different things, and a bare 404 makes an operator go and read
|
|
the mesh to tell them apart.
|
|
|
|
*Checked in the lab by a request to the name reaching the workload across the private network and
|
|
returning the workload's own answer, then by unassigning the module and requiring the same request
|
|
to stop working.*
|
|
|
|
### The name is a label, not a domain
|
|
|
|
*2026-09-09, from running the whole mesh in the lab.* A route contribution carried its public name
|
|
in full — the forge as `git.example.tld`, spelt out in the module. Pointing the same catalogue at a
|
|
different domain — a lab standing in for production, or a second operator's mesh — meant rewriting
|
|
that name on every routed module. The mesh was holding a **map of names to services**, which is the
|
|
one thing it must not: a public name is two facts owned by two different places, and neither is the
|
|
module's manifest.
|
|
|
|
**A module contributes a label; the node contributes its public domain; the mesh composes.** The
|
|
operator chooses where the forge lives — `git`, or `code` — and that is the module's to say. The
|
|
domain is the node's, set once. The mesh joins them and grants `<label>.<public-domain>`,
|
|
interpreting neither half. Moving a mesh to another domain is one node setting, not an edit per
|
|
module.
|
|
|
|
**Today the name is still a literal, and that is the gap.** There is no interpolation of a node's
|
|
domain into a module's label, so the composition is done by a per-node override — which reproduces
|
|
exactly the per-module cost it is meant to remove. The design is the composition; the override is a
|
|
stopgap until the manifest layer can carry a label and a domain separately
|
|
([ADR 0066](../../02-DECISIONS/0066-public-routing-is-name-agnostic.md)).
|
|
|
|
## 4 — Filtering
|
|
|
|
**Derived from what is assigned here, and from the overlay's shape** — a node's open ports are a
|
|
consequence of what runs on it and who must reach it, not an independent declaration to keep in
|
|
step by hand.
|
|
|
|
**A rule names its source** ([ADR 0007](../../02-DECISIONS/0007-connectivity.md)).
|
|
A rule with no source is open, and must say so rather than appear to restrict something. `scope:`
|
|
is removed rather than implemented: five manifests carry it today, it is referenced by no code,
|
|
and it is the clearest instance in the repository of *an unenforced rule is indistinguishable
|
|
from a wrong one, and costs more, because people believe it.*
|
|
|
|
**Unknown keys are refused** — the discipline the host's declaration parser already has
|
|
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)), and
|
|
the one manifests lack. `scope:` survived because nothing rejected it.
|
|
|
|
### What was built
|
|
|
|
*2026-08-31. Everything above was the intention; this is what exists, and how each part is
|
|
checked. [04-ISSUES/003](../../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) is
|
|
resolved by it.*
|
|
|
|
**A module says what it listens on**, as a port, a protocol and a source — `mesh`, `anywhere`, or
|
|
`machine`. The source is required and there is no default, which is the whole of *a rule names its
|
|
source*: a manifest that omitted it would read as a restriction and be none. *Checked by a manifest
|
|
with a port and no source being refused, and by one naming a source the mesh cannot render being
|
|
refused as well — the second is what stops a source becoming a comment.*
|
|
|
|
**The set is derived per node**, from every module assigned to it, not from the module asking for
|
|
it. Where two modules want the same port, the wider source wins and both are still named, because
|
|
removing one of them must not read as a reason to close a port the other needs. *Checked by
|
|
rendering a node whose firewall module has no ports of its own and asserting another module's port
|
|
is in the result; and by giving one port two modules and one source each, and asserting the
|
|
narrower rule disappears while both names survive.*
|
|
|
|
**What is not declared is closed.** The rule set drops by default. *Checked by naming the input
|
|
chain in the assertion rather than the policy alone — the first version of that test passed while
|
|
input accepted everything, because another chain in the same file also said `policy drop`.*
|
|
|
|
**From the mesh means the machines the mesh has**, as their addresses on the private network, not
|
|
as a subnet. A subnet is a guess that stays wrong quietly; the address set shrinks when a node
|
|
leaves and nobody edits anything. A machine that asks for `mesh` where the mesh knows no addresses
|
|
is **closed and told so in the file** — widening it would open a port nobody asked to open, and
|
|
dropping it silently would close one somebody did.
|
|
|
|
**Three things it deliberately does not do**, each of which looked right and would have broken
|
|
something:
|
|
|
|
| | why not |
|
|
|---|---|
|
|
| decide what the machine **forwards** | the container runtime writes its own forwarding rules and a second policy is consulted alongside them, so a drop here stops every container on the node — the controller included. Nothing in a manifest says what a machine routes, so there is nothing to derive it from either |
|
|
| **flush the ruleset** when loading | that empties every table on the machine, the runtime's among them. Only the mesh's own table is replaced, and it is declared empty first so the replacement works on a machine loading one for the first time |
|
|
| carry a **command to load itself** | the link may not carry an action ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)). A service is declared to reflect the file instead, so replacing it restarts what loads it — the shape that rule leaves, used here for the first time for its real purpose |
|
|
|
|
**A module that wants a rule set brings the unit that loads it.** Found the hard way: the unit a
|
|
distribution packages for nftables runs, applies the rules and exits, so it is neither running nor
|
|
stopped — and a host asked for a service that is "running" reports, quite correctly, that it is
|
|
stopped. Every packet was filtered exactly as declared and the machine was marked as not doing what
|
|
it was told.
|
|
|
|
**The vocabulary has no word for "ran, did its job, and exited"**, and that is a real gap rather
|
|
than a wording problem: the whole class of configuration-applying units — packet filters, sysctl,
|
|
tmpfiles — is shaped that way. Until there is one, a module ships a unit that stays, which is also
|
|
the better shape: how a machine enforces rules is a fact about the machine, and the mesh has no
|
|
business depending on what a distribution happens to package.
|
|
|
|
**One rule is derived from the overlay's shape rather than from what is assigned: a hub's own
|
|
listening port.** A hub accepts inbound connections from every node at other sites; a machine that
|
|
is not a hub dials out and needs nothing open, because a reply to a flow it started is already
|
|
accepted. The two want different rules on an *identical module*, so `listens` — a static field —
|
|
cannot say it. The machine a static answer gets wrong is the one facing the public internet, which
|
|
is the machine that most needs filtering.
|
|
|
|
*Recorded as a gap on 2026-08-31 and closed the same day.* **A computed module now contributes
|
|
listens the way it contributes resources.** The port comes from the endpoint, which is where the
|
|
interface takes its `ListenPort` from — one source, so a rule set cannot open a port the interface
|
|
is not on. It is open to *everywhere* deliberately: a node at another site is not on the private
|
|
network until this port lets it on, so restricting it to the mesh would be a rule that can never
|
|
be satisfied by the thing it exists for.
|
|
|
|
**A generator that cannot say what a machine opens is refused, not read as silence.** Closing a
|
|
port on the evidence of a failure to look is how a machine is severed by a fault somewhere else —
|
|
and the machine it would sever is the hub, whose only route to being repaired is the network it
|
|
just closed.
|
|
|
|
*Checked by filtering the hub and then requiring the mesh to keep working: a declaration still
|
|
reaches the other machine, and the other machine still reaches the hub. A rule file that looks
|
|
right and a mesh that has stopped are exactly what that guards against.*
|
|
|
|
**And it is enforced, which is what separates this from `scope:`.** Checked on two real machines:
|
|
two ports opened, one declared, and from the other machine the declared one answers and the
|
|
undeclared one does not — then the module is removed and the port closes with nobody editing a
|
|
rule. *A rule set that is written but never loaded passes every check that reads the file, which
|
|
is why the check reads packets.*
|
|
|
|
## 5 — Certificates
|
|
|
|
**Two authorities, kept separate on purpose.**
|
|
|
|
| | issued by | for |
|
|
|---|---|---|
|
|
| **public names** | a public ACME authority | anything outside the mesh reaches |
|
|
| **internal names** | the **mesh CA** | node-to-node, over the overlay |
|
|
|
|
**The split is not collapsed, including in the lab.** A single-CA lab would hide any bug living
|
|
in the split, so the lab runs its own ACME issuer on its public segment and keeps the mesh CA
|
|
unchanged ([research 004](../../01-RESEARCH/004-lab-network/00-overview.md)).
|
|
|
|
**Public issuance requires genuine public reachability.** The HTTP-01 challenge must be answered
|
|
at the name being certified, so issuance happens through a publicly reachable node regardless of
|
|
where the workload runs — the same asymmetry as exposure, for the same reason.
|
|
|
|
**The issuer must be configurable.** Today it is not: the proxy sets no `caServer` and therefore
|
|
defaults to the public authority's *production* endpoint. Two consequences, and the second is
|
|
worse than the lab problem that found it — every certificate experiment on a real node consumes
|
|
production issuance quota, and a retry loop can exhaust it for a week.
|
|
|
|
**The mesh CA is not a bootstrap concern.** A joining node verifies the controller against the
|
|
fingerprint in its token ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)),
|
|
so nothing needs the CA before membership. It certifies internal names afterwards, and that is
|
|
all it does.
|
|
|
|
### What was built
|
|
|
|
*2026-08-31.*
|
|
|
|
**A node generates a fourth key**, and the reason is the one the other three already give: *a key
|
|
used for two purposes is one rotation away from breaking the other.* The identity key would work
|
|
for TLS and reusing it would mean rotating a node's identity every time its certificate is
|
|
replaced. The private half never leaves the machine; the mesh is told the public half at
|
|
enrolment.
|
|
|
|
**So there is no certificate request and nothing to seal.** The mesh signs a statement binding a
|
|
public key to a name it alone assigns, which is the whole of what a certificate authority does.
|
|
It issues rather than stores: the node's key does not change, so signing again produces an equally
|
|
valid certificate and there is nothing to keep in step.
|
|
|
|
**A machine with no name inside the mesh is refused**, not given a certificate for nothing. A
|
|
certificate for a name nothing resolves is a certificate nothing can check.
|
|
|
|
**And the key is stored in the format a server reads** — PKCS#8 PEM, not the host's own encoding.
|
|
That is not an implementation detail of whoever writes the file: the file exists *because
|
|
something else reads it*, so the format is the interface
|
|
([04-ISSUES/014](../../04-ISSUES/014-a-key-that-is-present-and-unusable/00-report.md)).
|
|
|
|
*Checked by a real handshake between two machines: one serves on its internal name with the key it
|
|
generated, the other verifies against the mesh's authority and nothing else. Every cheaper check
|
|
passed while the server could not start — the key was present, the certificate was valid, and
|
|
nothing read either the way a server would.*
|
|
|
|
### An internal issuer, pointed at and trusted
|
|
|
|
*2026-09-09, from wiring one to the proxy in the lab.* §5 above asks for two things this build
|
|
leaned on: the issuer must be configurable, and the lab runs its own ACME authority rather than
|
|
collapsing the split. Wiring the proxy to that authority is the whole of it — the proxy is told
|
|
**which** issuer to use and given that issuer's **root** to trust, and every other step of issuance
|
|
is unchanged. **The same code path certifies against an internal authority as against a public one;
|
|
only the issuer differs.** That is what makes trusted certificates possible for a mesh whose names
|
|
the public internet cannot resolve.
|
|
|
|
**And it does not work until the routed name resolves inside the mesh** — the §2 finding above,
|
|
arriving here because this is what needed it. The authority's challenge reaches the routed name only
|
|
once that name is in internal resolution; a public authority is handed that dependency by public
|
|
DNS, and an internal one has to be handed it by the mesh. *Checked by a handshake to a routed name
|
|
that verifies against the internal root and nothing else — which cannot succeed unless the issuer
|
|
first reached the name to certify it*
|
|
([ADR 0066](../../02-DECISIONS/0066-public-routing-is-name-agnostic.md)).
|
|
|
|
## What this removes
|
|
|
|
The list is worth having in one place, because it is most of the argument:
|
|
|
|
- **The last two direct database connections from nodes** — `wireguard` and `traefik`, the only
|
|
two, both connectivity.
|
|
- **Therefore the database credential on every node**, and the object-store credential beside it.
|
|
[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)'s central claim becomes
|
|
true rather than aspirational.
|
|
- **The `/etc/hosts` floor**, and the bootstrap circularity it patched.
|
|
- **Hub election by address prefix**, and the silent no-hub failure when nobody knew the
|
|
convention.
|
|
- **The RFC1918 inference**, and the lab substitution that existed to satisfy it.
|
|
- **`scope:`**, and the class of manifest key that means nothing.
|
|
|
|
## Open
|
|
|
|
- ~~**What happens when the hub is down.**~~ **Resolved** by
|
|
[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md), together with `06`'s
|
|
matching question — they were one question. Nothing takes over. WireGuard has no failover, the
|
|
hub is declared rather than elected, and non-co-located paths stop while co-located direct peers
|
|
and every already-assigned workload keep running. The recovery path is restore, and its deadline
|
|
is certificate renewal.
|
|
- **Renumbering the overlay.** Made *possible* by declaring the hub rather than inferring it from
|
|
an address, but no procedure exists, and a graph delivered node by node has an ordering problem
|
|
while it is half-applied.
|
|
- ~~**Revoking a route** when a module is unassigned.~~ **Resolved** 2026-08-31 — see §3. The file
|
|
a proxy is given is the whole truth about who has a route, so a route does not outlive the module
|
|
that asked for it.
|
|
- **IPv6.** [ADR 0007](../../02-DECISIONS/0007-connectivity.md) makes
|
|
it expressible; nothing here says the overlay or the resolver handle it.
|
|
- **Reporting declared-versus-observed.** ADR 0007 makes the disagreement detectable and does not
|
|
say who looks or what they are told.
|
|
- **Composing a route name from a label and a node's domain.**
|
|
[ADR 0066](../../02-DECISIONS/0066-public-routing-is-name-agnostic.md) makes the public name the
|
|
operator's to move between meshes, but the manifest layer still stores it as a literal — so today
|
|
the composition is a per-node override rather than the design. The interpolation that would let a
|
|
module carry a label and a node carry the domain, and the mesh join them, does not yet exist.
|
|
- **Publishing route names into internal resolution.** The same ADR requires a granted route to be
|
|
resolvable inside the mesh, not only routable from outside it; the mechanism that writes
|
|
`<node>.internal` into containers does not yet also write the routed names, which is why an
|
|
internal issuer cannot currently validate one without a hand-placed entry.
|