8b2d1bce9dd4e1edea4a525f3cb15338977e9150
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
44d134ba25 |
Networking is a module, and a domain module is how you avoid choosing
Connectivity was code beside the module system doing the module system's
job: every machine with an address was on the private network and there
was no way to keep one off.
A manifest can now say its resources are computed by the control plane,
which is what a peer list needs — it is derived from every machine at
once, so nothing could be written in advance. The network is a module
from there on: assigned, resolved, settled, and absent from a machine
nobody gave it to.
Three modules rather than one, because WireGuard is one VPN of several:
mesh-wireguard provides private-network, mesh-addressing
claims the-private-network, one per node
mesh-names provides name-resolution, requires mesh-addressing
networking requires both, and ships no files of its own
The last is the point. Most people want the network up and do not want
to choose a VPN, so `assign networking` takes the only answer to each
requirement silently. The day the catalogue holds a second one there are
two answers, the resolver refuses and names them, and choosing is
assigning the one you want. No flavor field, nothing to configure.
Names left the WireGuard declaration for their own module. They would be
identical over a different private network, and bundling them made one
module out of two things.
Three faults the walk found:
- choosing tailscale still installed WireGuard, dragged back in by the
names needing the mesh's own addresses. Caught now by a claim: running
two VPNs is fine, being *the* mesh network is singular.
- a requirement wanted by two modules was reported twice, identically.
- "this mesh has no hub" was reported when the real cause was that a
node could not be resolved at all. It now names the node and the why.
And a test that asserts the manifests actually shipped, after the claim
went missing from the real one while every test stayed green.
|
||
|
|
fc1417be72 |
Names, from the same graph as the network
Step 5 of the connectivity order. Every node's internal name resolves to its overlay address, on every node, computed centrally because it needs every node at once. Under `.internal`, which IANA reserved for exactly this in 2024 -- a name there can never collide with a public one, so an internal name that leaks into a public resolver fails rather than reaching a stranger's machine. The suffix is settable for a mesh that wants its own. Delivered in the same declaration as the peer list rather than a second one. A node holding the peers and not the names, or the reverse, is half on the network for as long as that lasts. This is not the /etc/hosts floor the design removes. That floor existed because a node had to reach the mesh's database before its own DNS worked -- a fallback for a circularity that is now gone. This is the mechanism: the complete set of names, generated whole and owned by the mesh, rather than a patch written underneath something else. A resolver daemon becomes necessary when names are wanted that are not one-per-node, and that is not yet true. A node resolves its own name to its overlay address rather than a loopback, because a service binding to the name it was given would otherwise listen somewhere nothing else can reach -- and the failure would appear on every other machine rather than that one. A node with no address gets no name. A name resolving to nothing is worse than no name: connecting to an address that does not answer hangs, where a name that does not resolve fails at once and says which name it was. Found while writing it: a test asserting every file in the declaration is mode 0600 would have forced /etc/hosts to 0600 and broken every lookup on the machine, to protect a file that is not secret. Verified in the lab: three machines, nine name lookups, each resolving to the right overlay address and reaching it. |
||
|
|
8b974deb42 |
A working private network, and four reasons it did not work
Three machines across two sites, two of them behind no reachable address, all nine paths open. The mesh computes the graph, delivers it as a declaration, and the nodes bring it up. Every fault below looked like success from inside the mesh: the graph was right, the files were right, the services were up, every node reported it had applied. None was reachable by reasoning. A running interface does not re-read its configuration. A node joins, every existing node's peer list changes, the file is replaced -- and the service is already running, so nothing reloads it. Fixed as declared state rather than a command: the service must reflect the file. A command to restart would be an action, and the link may not carry one. The host refused exactly that, which is how this shape was arrived at. A hub sharing a site with a spoke appeared twice in that spoke's peer list -- once as a direct peer, once as the route of last resort. WireGuard takes one entry per key and refuses the file. The ordinary shape of a small mesh, and in none of the tests written before it ran. Two nodes at one site that neither can be dialled were peered directly. Nobody opens the path, and the direct route is more specific than the hub's, so it wins and blackholes -- this design's own warning arriving in its implementation. They now route through the hub unless one end can be dialled. And Docker sets the FORWARD policy to DROP, so a hub with ip_forward enabled carried nothing between its spokes. The substrate at tier 1 silently breaks the network at tier 2, and nothing in either tier's state says so. The hub inserts its own rule above those chains and removes it on the way down. Two weak tests found by injection along the way: one asserted the keepalive rule only against the hub, whose peer entries happen not to set that field at all, so it tested an absence; the other checked the firewall rules by looking for FORWARD anywhere, which the PostDown line satisfies on its own. |
||
|
|
f44e73d286 |
The mesh computes a private network it cannot impersonate
The first thing the control plane decides rather than relays. Every node's peer list is derived from every node at once, which is what makes this control-plane work by definition: no node has that view. A hub, with direct peering between nodes at the same site. Not a full mesh, and the reason is a property of WireGuard rather than a preference -- there is no failover, so a more specific route to a dead endpoint blackholes instead of falling back. A node gets exactly one path to any peer, because two would mean one of them silently swallowing traffic. A roaming node is hub-only for the same reason. Reachability and the hub are declared, never inferred from an address. The address is evidence and is not the fact: carrier-grade NAT looks public and is not, a routable address behind a closed firewall looks public and is not, and the regular expression that used to decide it got the lab wrong too. Hub election by address prefix failed silently when nobody knew the convention. No private key travels, and that is the whole design. The node generated its own keypair and kept the private half; the configuration points at a file the node wrote, using WireGuard's own PostUp. So the control plane composes a complete configuration for a node it cannot pretend to be -- it knows every public key and holds none of the private ones. Delivered as an ordinary declaration: a package, a file and a service. The host does not know what a private network is and does not learn one. There is a test holding that line, because the moment connectivity needs a new shape in tier 0 is the moment the host stops being small enough to trust. The generated file is written to be read: each peer says why it is there, a peer with no endpoint says why it has none, and the header says not to edit it -- an edit survives until the graph next changes and then vanishes, which is worse than never being applied, because the machine works and then stops and nothing changed that anybody remembers. Fault injection found one weak test. The keepalive rule was asserted only against the hub, whose peer entries happen not to set the field at all, so it was testing an absence rather than the rule. It now checks two direct peers where one is reachable and one is not. |