A working private network, and four reasons it did not work
Three machines across two sites, two of them behind no reachable address, all nine paths open. The mesh computes the graph, delivers it as a declaration, and the nodes bring it up. Every fault below looked like success from inside the mesh: the graph was right, the files were right, the services were up, every node reported it had applied. None was reachable by reasoning. A running interface does not re-read its configuration. A node joins, every existing node's peer list changes, the file is replaced -- and the service is already running, so nothing reloads it. Fixed as declared state rather than a command: the service must reflect the file. A command to restart would be an action, and the link may not carry one. The host refused exactly that, which is how this shape was arrived at. A hub sharing a site with a spoke appeared twice in that spoke's peer list -- once as a direct peer, once as the route of last resort. WireGuard takes one entry per key and refuses the file. The ordinary shape of a small mesh, and in none of the tests written before it ran. Two nodes at one site that neither can be dialled were peered directly. Nobody opens the path, and the direct route is more specific than the hub's, so it wins and blackholes -- this design's own warning arriving in its implementation. They now route through the hub unless one end can be dialled. And Docker sets the FORWARD policy to DROP, so a hub with ip_forward enabled carried nothing between its spokes. The substrate at tier 1 silently breaks the network at tier 2, and nothing in either tier's state says so. The hub inserts its own rule above those chains and removes it on the way down. Two weak tests found by injection along the way: one asserted the keepalive rule only against the hub, whose peer entries happen not to set that field at all, so it tested an absence; the other checked the firewall rules by looking for FORWARD anywhere, which the PostDown line satisfies on its own.
This commit is contained in:
@@ -62,6 +62,19 @@ func Declaration(node Node, peers []Peer, keyPath string) ([]byte, error) {
|
||||
// be told again. A node whose overlay only exists while something is watching is not
|
||||
// a node that survives being switched off and on.
|
||||
"boot": "enabled",
|
||||
// And restarted when the peer list changes, because a running interface does not
|
||||
// re-read its configuration.
|
||||
//
|
||||
// This is the whole of it: a node joins, every existing node's peer list changes,
|
||||
// each file is replaced — and without this the service is already running, nothing
|
||||
// reloads it, and every node keeps a network that no longer matches the mesh. It
|
||||
// reports complete success. The lab found it the moment a third node arrived.
|
||||
//
|
||||
// Declared state rather than a command. The service must reflect the file; the host
|
||||
// works out that it does not. A command to restart would be an action, and the link
|
||||
// may not carry one (novox/hq ADR 0005) — the host refused exactly that, correctly,
|
||||
// which is how this shape was arrived at.
|
||||
"restart-on": []string{"overlay-config"},
|
||||
},
|
||||
}
|
||||
|
||||
@@ -96,6 +109,39 @@ func config(node Node, peers []Peer, keyPath string) string {
|
||||
// travelled. Everything else in this file came from the mesh; this one line is the node's.
|
||||
fmt.Fprintf(&b, "PostUp = wg set %%i private-key %s\n", keyPath)
|
||||
|
||||
if node.Hub {
|
||||
// A hub carries traffic *between* its spokes, and a Linux machine does not forward
|
||||
// packets unless it is told to. Without this every spoke reaches the hub perfectly and
|
||||
// no spoke reaches any other — which is exactly how it failed in the lab, and it looks
|
||||
// like a peering problem rather than a kernel setting.
|
||||
//
|
||||
// Here rather than in a separate resource because it is part of what being a hub means,
|
||||
// and because it should last exactly as long as the interface does: a machine that stops
|
||||
// being the hub should stop forwarding, and PostDown below is how that happens.
|
||||
b.WriteString("PostUp = sysctl -q -w net.ipv4.ip_forward=1\n")
|
||||
b.WriteString("PostDown = sysctl -q -w net.ipv4.ip_forward=0\n")
|
||||
|
||||
// And past the machine's own firewall, which on any node with a container runtime is
|
||||
// closed. Docker sets the FORWARD policy to DROP and inserts its chains, so the substrate
|
||||
// this mesh installs at tier 1 silently breaks the network it builds at tier 2: every
|
||||
// spoke reaches the hub, no spoke reaches any other, and every part of it reports
|
||||
// success. Found in the lab; nothing about it is visible from the mesh's own state.
|
||||
//
|
||||
// Inserted at the top so it precedes those chains, and removed on the way down so a
|
||||
// machine that stops being the hub stops carrying other people's traffic. Guarded on
|
||||
// iptables existing: a machine without it has no policy to get past.
|
||||
for _, direction := range []string{"-i", "-o"} {
|
||||
fmt.Fprintf(&b,
|
||||
"PostUp = command -v iptables >/dev/null && iptables -I FORWARD 1 %s %%i -j ACCEPT || true\n",
|
||||
direction)
|
||||
}
|
||||
for _, direction := range []string{"-i", "-o"} {
|
||||
fmt.Fprintf(&b,
|
||||
"PostDown = command -v iptables >/dev/null && iptables -D FORWARD %s %%i -j ACCEPT || true\n",
|
||||
direction)
|
||||
}
|
||||
}
|
||||
|
||||
for _, p := range peers {
|
||||
fmt.Fprintf(&b, "\n# %s — %s\n[Peer]\n", p.Name, p.Why)
|
||||
fmt.Fprintf(&b, "PublicKey = %s\n", p.Key)
|
||||
|
||||
Reference in New Issue
Block a user