Files
mesh-controller/lab/adopt-the-tunnel/README.md
T
jschoubben 3c836f0abb Adopt the predecessor's tunnel in place: its range, its address, its peers
On an adopted hub the private network takes over the tunnel it finds rather
than running beside it (hq ADR 0105): two tunnels leave the mesh's unreachable
through the provider's filter, so no machine can ever join.

The node presents the found tunnel when it enrols, under the key it took as
its own; the inventory records it (node.tunnel, tunnel_peer — migration 0031)
and the mesh composes from it: the overlay's range is the adopted tunnel's,
the hub is placed at the tunnel's address on the tunnel's port, and every
peer the tunnel had is carried in the hub's peer list as a peer of the
tunnel, not a node of the mesh, until a node enrols with that key — which
then keeps the address the tunnel had for it. A fresh node never gets an
address the tunnel holds. The hub's declaration tells the host which unit to
take over; the host's account of carrying it is recorded and shown.

Every reader of the range follows the setting; nothing stores it. A found
tunnel under another key is recorded and not adopted, so ADR 0100's
non-overlap rule keeps applying where a tunnel is left running beside the
mesh's. A lab bed and test skeleton for "How it is checked" are under lab/.
2026-09-23 23:26:34 +02:00

6.2 KiB

Lab bed: the hub adopts the predecessor's tunnel (novox/hq ADR 0105)

A scenario and an integration-test skeleton for the mesh-lab repository, kept here because this branch changes only the controller and the host. Move adopt-the-tunnel.yml to mesh-lab/scenarios/ and adopt-the-tunnel.test.ts to mesh-lab/test/integration/ when the feature lands; neither has been run. The skeleton follows adoption.test.ts and reuses its harness. Documentation addresses throughout; the bed is node-agnostic.

The bed, precisely

Three machines on one public segment, hosting (192.0.2.0/24), inbound allowed on all (the anchor's firewall is the predecessor's, installed by the bed):

machine address role
anchor 192.0.2.10 the machine in use: the predecessor's hub, then the mesh adopted on it
peer-a 192.0.2.20 a predecessor machine: reaches a service on the anchor through the tunnel; later enrols and keeps its address
peer-b 192.0.2.30 a second predecessor machine: reaches the same service; never enrols — the peer that must notice nothing throughout
fresh 192.0.2.40 a new machine: enrols later and gets a fresh address from the same range

Prepared the way the predecessor leaves a hub (before genesis, by the bed, on anchor):

  • wireguard-tools installed; a keypair made on each of anchor, peer-a, peer-b.
  • /etc/wireguard/wg0.conf on the anchor: [Interface] PrivateKey = <anchor's>, ListenPort = 51900, Address = 10.10.0.1/24; two [Peer] sections — peer-a's public key with AllowedIPs = 10.10.0.2/32, peer-b's with AllowedIPs = 10.10.0.3/32. Raised with systemctl enable --now wg-quick@wg0. 10.10.0.0/24 is deliberately not the mesh's default range (10.42.0.0/16), so a hub address in 10.10.0.0/24 can only have come from the tunnel.
  • wg0.conf on each peer: its own key, Address = 10.10.0.2/24 (resp. .3/24), one [Peer] — the anchor's public key, Endpoint = 192.0.2.10:51900, AllowedIPs = 10.10.0.0/24, PersistentKeepalive = 25. Raised the same way.
  • A service on the anchor the peers reach only over the tunnel: a container publishing 10.10.0.1:8081:80 (bound to the tunnel address, so a call from 192.0.2.20 to 10.10.0.1:8081 proves the tunnel carried it). Under a name no catalogue module uses — this bed is about the tunnel, not about taking a service.
  • The predecessor's firewall (ufw) allowing 51900/udp and 22/tcp, denying the rest — as ADR 0100's bed prepares it.
  • A record of the anchor's wg0 public key and of sha256sum /etc/wireguard/wg0.conf, taken before genesis, for the assertions below.

Genesis, adopted, on the anchor: mesh-bootstrap --adopted --node anchor --site hosting --endpoint 192.0.2.10:51900 … — no --hub-port, no --overlay-range, no --tunnel: the installer finds wg0 itself (one interface besides mesh0) and takes its port and range. The bed asserts genesis says it found and took the tunnel.

Assertions, in the record's order

  • T1 — the tunnel changes hands and the peers notice nothing. After genesis and the push that raises the private network on the anchor:
    • wg show interfaces on the anchor lists mesh0 and not wg0; systemctl is-active wg-quick@wg0 is inactive and is-enabled disabled;
    • /etc/wireguard/wg0.conf is on disk with the recorded digest (kept, never flushed), and node show anchor names where its original was kept;
    • wg show mesh0 public-key is the anchor's recorded wg0 public key; wg show mesh0 listen-port is 51900; ip -o addr show dev mesh0 carries 10.10.0.1; wg show mesh0 peers lists both peers' public keys with their /32 allowed addresses;
    • a loop on peer-a and peer-b calling http://10.10.0.1:8081/ every second, started before genesis, records no window of failure longer than one WireGuard re-handshake (measure and assert an upper bound — the switch is one unit stop plus one unit start on the anchor); the peers' wg0.conf digests are unchanged; the peers' wg show wg0 latest-handshakes advance after the switch.
    • overlay show lists anchor as the hub over the tunnel it took over, and both peers under "peers of the tunnel … not nodes of the mesh", not yet enrolled.
  • T2 — a peer enrols and keeps its address. On peer-a: node add peer-a --adopted, token issued, mesh-host enrol --token … with the broker reached over the tunnel (the broker address in the token is 10.10.0.1:<bus>, which only the tunnel routes); then overlay place peer-a --site house and a push. Assert: node show peer-a says a tunnel wg0 was found and overlay show puts peer-a at 10.10.0.2; the carried-peers list now says enrolled as peer-a; the anchor's mesh0 still has exactly one entry for peer-a's key; peer-a's wg0 is down and mesh0 up with the same key; peer-a still reaches 10.10.0.1:8081 and peer-b still does too, uninterrupted.
  • T3 — a new machine gets a fresh address from the same range. fresh enrols converged (no tunnel), is placed, pushed. Assert its address is 10.10.0.4 (.1 hub, .2 and .3 the tunnel's), that it reaches 10.10.0.1 (the hub) and 10.10.0.2 (the enrolled peer) — ping -c1 over mesh0 — and that peer-b, never enrolled, is still served.
  • T4 — nothing derived from the address is stale. plan anchor --json and plan peer-a --json (and the rendered /etc/hosts on each node) name 10.10.0.1 for the anchor and 10.10.0.2 for peer-a, and no address in 10.42.0.0/16; the broker address handed to a module issued on peer-a is anchor.internal:<bus> resolving to 10.10.0.1; the same after a second push of every node, byte for byte.
  • N — the narrowed ADR 0100 check. On fresh, a converged genesis dry-run with --overlay-range 10.10.0.0/24 while a second tunnel the bed raises there (wg1 at 10.10.0.9/24, not adopted because the node is converged) is up, is refused naming wg1 — the non-overlap rule still applies where a tunnel is not adopted.

What is not asserted here

  • Taking a service over the tunnel (ADR 0100's bed does that).
  • IPv6 tunnels: the parser reads them, the bed prepares only IPv4.