Commit Graph
6 Commits
Author SHA1 Message Date
jschoubben 7fc5fd02fd The mesh's interface takes over the found tunnel's MTU
Carries MTU from the reported tunnel (mesh-host#28) through inventory,
the overlay graph's TakeOver, into the generated config's [Interface].
A tuned path keeps its MTU across the takeover instead of regressing to
1420 and hanging transfers no ping would reveal. Two emit tests; a
tunnel with no MTU writes no line.
2026-09-26 22:40:42 +02:00
jschoubben cc252472e2 A taken tunnel brings its ListenPort, even on a node the hub cannot dial
A home node behind NAT (no Endpoint → not Reachable) that took over a
tunnel must still listen on that tunnel's port: its LAN peers dial it
there. ListenPort was gated on Reachable, which conflated 'a peer dials
me here' with 'the hub can dial me' — so the takeover guard refused
overlay-up, and the guard's suggested remedy (re-place with an
endpoint) breaks a NAT'd node's path: it stops keepalive and hands the
hub a private LAN address to dial. TakeOver now carries the found
tunnel's port (already known to the controller), and the interface
listens on it when the node is not otherwise reachable. Two tests;
Endpoint-reachable nodes keep the old path unchanged.
2026-09-26 22:33:27 +02:00
jschoubben 4566c5c9aa Adopt the tunnel as a mesh fact, refuse a mismatched takeover, and rekey after enrolment
Review of the ADR 0105 build (hq ADR 0105). Four things it got wrong and one
path it lacked:

- A predecessor spoke's tunnel names one peer, the hub, routed the whole
  range; recording refused it and the whole enrolment failed. Range-routed
  peers are skipped now — only the hub's peers are ever carried.
- The range and the carried peers were conditions on the node being adopted,
  so converging the hub would have renumbered the mesh and dropped the peers
  still reaching it. They are facts of the tunnel record now, mode aside; the
  takeover alone is declared to an adopted node. Converging the hub is refused
  while a carried peer has not enrolled, naming it.
- A push composed a takeover for a hub whose address or endpoint disagreed
  with the tunnel, which would have the host stop the found interface and
  raise the mesh's where no peer listens. The graph refuses to compose it,
  naming both and the placement that fixes it.
- The host's account said taken or not; "found down and the mesh's not up"
  read as not taken. Three states now, and an account on every takeover.
- A hub that enrolled before this feature holds a key of its own, and
  re-enrolling would rotate every key the mesh sealed credentials to. A node
  now rekeys in a report, signed with its identity key over the key it
  leaves, the key it takes and the tunnel; the mesh verifies against the live
  key, refuses a stale or foreign proof, records key and tunnel, and moves a
  hub to the tunnel's address. `overlay show` names the path for a hub that
  found no tunnel.

Also: a carried IPv6 peer is routed /128, and identity.ForTest exists so the
link can be tested against a real identity store.
2026-09-24 00:02:07 +02:00
jschoubben 3c836f0abb Adopt the predecessor's tunnel in place: its range, its address, its peers
On an adopted hub the private network takes over the tunnel it finds rather
than running beside it (hq ADR 0105): two tunnels leave the mesh's unreachable
through the provider's filter, so no machine can ever join.

The node presents the found tunnel when it enrols, under the key it took as
its own; the inventory records it (node.tunnel, tunnel_peer — migration 0031)
and the mesh composes from it: the overlay's range is the adopted tunnel's,
the hub is placed at the tunnel's address on the tunnel's port, and every
peer the tunnel had is carried in the hub's peer list as a peer of the
tunnel, not a node of the mesh, until a node enrols with that key — which
then keeps the address the tunnel had for it. A fresh node never gets an
address the tunnel holds. The hub's declaration tells the host which unit to
take over; the host's account of carrying it is recorded and shown.

Every reader of the range follows the setting; nothing stores it. A found
tunnel under another key is recorded and not adopted, so ADR 0100's
non-overlap rule keeps applying where a tunnel is left running beside the
mesh's. A lab bed and test skeleton for "How it is checked" are under lab/.
2026-09-23 23:26:34 +02:00
jschoubben 8b974deb42 A working private network, and four reasons it did not work
Three machines across two sites, two of them behind no reachable address, all
nine paths open. The mesh computes the graph, delivers it as a declaration, and
the nodes bring it up.

Every fault below looked like success from inside the mesh: the graph was
right, the files were right, the services were up, every node reported it had
applied. None was reachable by reasoning.

A running interface does not re-read its configuration. A node joins, every
existing node's peer list changes, the file is replaced -- and the service is
already running, so nothing reloads it. Fixed as declared state rather than a
command: the service must reflect the file. A command to restart would be an
action, and the link may not carry one. The host refused exactly that, which is
how this shape was arrived at.

A hub sharing a site with a spoke appeared twice in that spoke's peer list --
once as a direct peer, once as the route of last resort. WireGuard takes one
entry per key and refuses the file. The ordinary shape of a small mesh, and in
none of the tests written before it ran.

Two nodes at one site that neither can be dialled were peered directly. Nobody
opens the path, and the direct route is more specific than the hub's, so it
wins and blackholes -- this design's own warning arriving in its
implementation. They now route through the hub unless one end can be dialled.

And Docker sets the FORWARD policy to DROP, so a hub with ip_forward enabled
carried nothing between its spokes. The substrate at tier 1 silently breaks the
network at tier 2, and nothing in either tier's state says so. The hub inserts
its own rule above those chains and removes it on the way down.

Two weak tests found by injection along the way: one asserted the keepalive
rule only against the hub, whose peer entries happen not to set that field at
all, so it tested an absence; the other checked the firewall rules by looking
for FORWARD anywhere, which the PostDown line satisfies on its own.
2026-08-29 18:04:15 +02:00
jschoubben f44e73d286 The mesh computes a private network it cannot impersonate
The first thing the control plane decides rather than relays. Every node's peer
list is derived from every node at once, which is what makes this control-plane
work by definition: no node has that view.

A hub, with direct peering between nodes at the same site. Not a full mesh, and
the reason is a property of WireGuard rather than a preference -- there is no
failover, so a more specific route to a dead endpoint blackholes instead of
falling back. A node gets exactly one path to any peer, because two would mean
one of them silently swallowing traffic. A roaming node is hub-only for the
same reason.

Reachability and the hub are declared, never inferred from an address. The
address is evidence and is not the fact: carrier-grade NAT looks public and is
not, a routable address behind a closed firewall looks public and is not, and
the regular expression that used to decide it got the lab wrong too. Hub
election by address prefix failed silently when nobody knew the convention.

No private key travels, and that is the whole design. The node generated its
own keypair and kept the private half; the configuration points at a file the
node wrote, using WireGuard's own PostUp. So the control plane composes a
complete configuration for a node it cannot pretend to be -- it knows every
public key and holds none of the private ones.

Delivered as an ordinary declaration: a package, a file and a service. The host
does not know what a private network is and does not learn one. There is a test
holding that line, because the moment connectivity needs a new shape in tier 0
is the moment the host stops being small enough to trust.

The generated file is written to be read: each peer says why it is there, a
peer with no endpoint says why it has none, and the header says not to edit it
-- an edit survives until the graph next changes and then vanishes, which is
worse than never being applied, because the machine works and then stops and
nothing changed that anybody remembers.

Fault injection found one weak test. The keepalive rule was asserted only
against the hub, whose peer entries happen not to set the field at all, so it
was testing an absence rather than the rule. It now checks two direct peers
where one is reachable and one is not.
2026-08-29 16:58:56 +02:00