On an adopted hub the private network takes over the tunnel it finds rather than running beside it (hq ADR 0105): two tunnels leave the mesh's unreachable through the provider's filter, so no machine can ever join. The node presents the found tunnel when it enrols, under the key it took as its own; the inventory records it (node.tunnel, tunnel_peer — migration 0031) and the mesh composes from it: the overlay's range is the adopted tunnel's, the hub is placed at the tunnel's address on the tunnel's port, and every peer the tunnel had is carried in the hub's peer list as a peer of the tunnel, not a node of the mesh, until a node enrols with that key — which then keeps the address the tunnel had for it. A fresh node never gets an address the tunnel holds. The hub's declaration tells the host which unit to take over; the host's account of carrying it is recorded and shown. Every reader of the range follows the setting; nothing stores it. A found tunnel under another key is recorded and not adopted, so ADR 0100's non-overlap rule keeps applying where a tunnel is left running beside the mesh's. A lab bed and test skeleton for "How it is checked" are under lab/.
90 lines
6.2 KiB
Markdown
90 lines
6.2 KiB
Markdown
# Lab bed: the hub adopts the predecessor's tunnel (novox/hq ADR 0105)
|
|
|
|
A scenario and an integration-test skeleton for the mesh-lab repository, kept here because this
|
|
branch changes only the controller and the host. Move `adopt-the-tunnel.yml` to
|
|
`mesh-lab/scenarios/` and `adopt-the-tunnel.test.ts` to `mesh-lab/test/integration/` when the
|
|
feature lands; neither has been run. The skeleton follows `adoption.test.ts` and reuses its
|
|
harness. Documentation addresses throughout; the bed is node-agnostic.
|
|
|
|
## The bed, precisely
|
|
|
|
Three machines on one public segment, `hosting` (192.0.2.0/24), inbound allowed on all (the
|
|
anchor's firewall is the predecessor's, installed by the bed):
|
|
|
|
| machine | address | role |
|
|
|---|---|---|
|
|
| `anchor` | 192.0.2.10 | the machine in use: the predecessor's hub, then the mesh adopted on it |
|
|
| `peer-a` | 192.0.2.20 | a predecessor machine: reaches a service on the anchor through the tunnel; later **enrols and keeps its address** |
|
|
| `peer-b` | 192.0.2.30 | a second predecessor machine: reaches the same service; **never enrols** — the peer that must notice nothing throughout |
|
|
| `fresh` | 192.0.2.40 | a new machine: enrols later and **gets a fresh address from the same range** |
|
|
|
|
**Prepared the way the predecessor leaves a hub** (before genesis, by the bed, on `anchor`):
|
|
|
|
- `wireguard-tools` installed; a keypair made on each of `anchor`, `peer-a`, `peer-b`.
|
|
- `/etc/wireguard/wg0.conf` on the anchor: `[Interface]` `PrivateKey = <anchor's>`,
|
|
`ListenPort = 51900`, `Address = 10.10.0.1/24`; two `[Peer]` sections — `peer-a`'s public
|
|
key with `AllowedIPs = 10.10.0.2/32`, `peer-b`'s with `AllowedIPs = 10.10.0.3/32`. Raised with
|
|
`systemctl enable --now wg-quick@wg0`. **10.10.0.0/24 is deliberately not the mesh's default
|
|
range** (10.42.0.0/16), so a hub address in 10.10.0.0/24 can only have come from the tunnel.
|
|
- `wg0.conf` on each peer: its own key, `Address = 10.10.0.2/24` (resp. `.3/24`), one
|
|
`[Peer]` — the anchor's public key, `Endpoint = 192.0.2.10:51900`,
|
|
`AllowedIPs = 10.10.0.0/24`, `PersistentKeepalive = 25`. Raised the same way.
|
|
- A service on the anchor the peers reach **only over the tunnel**: a container publishing
|
|
`10.10.0.1:8081:80` (bound to the tunnel address, so a call from 192.0.2.20 to 10.10.0.1:8081
|
|
proves the tunnel carried it). Under a name no catalogue module uses — this bed is about the
|
|
tunnel, not about taking a service.
|
|
- The predecessor's firewall (`ufw`) allowing `51900/udp` and `22/tcp`, denying the rest — as ADR
|
|
0100's bed prepares it.
|
|
- A record of the anchor's `wg0` public key and of `sha256sum /etc/wireguard/wg0.conf`, taken
|
|
before genesis, for the assertions below.
|
|
|
|
**Genesis**, adopted, on the anchor: `mesh-bootstrap --adopted --node anchor --site hosting
|
|
--endpoint 192.0.2.10:51900 …` — no `--hub-port`, no `--overlay-range`, no `--tunnel`: the
|
|
installer finds `wg0` itself (one interface besides `mesh0`) and takes its port and range. The
|
|
bed asserts genesis **says** it found and took the tunnel.
|
|
|
|
## Assertions, in the record's order
|
|
|
|
- **T1 — the tunnel changes hands and the peers notice nothing.** After genesis and the push
|
|
that raises the private network on the anchor:
|
|
- `wg show interfaces` on the anchor lists `mesh0` and not `wg0`;
|
|
`systemctl is-active wg-quick@wg0` is inactive and `is-enabled` disabled;
|
|
- `/etc/wireguard/wg0.conf` is on disk with the recorded digest (kept, never flushed), and
|
|
`node show anchor` names where its original was kept;
|
|
- `wg show mesh0 public-key` is the anchor's recorded `wg0` public key; `wg show mesh0
|
|
listen-port` is 51900; `ip -o addr show dev mesh0` carries `10.10.0.1`; `wg show mesh0 peers`
|
|
lists both peers' public keys with their `/32` allowed addresses;
|
|
- a loop on `peer-a` and `peer-b` calling `http://10.10.0.1:8081/` every second, started before
|
|
genesis, records **no window of failure longer than one WireGuard re-handshake** (measure and
|
|
assert an upper bound — the switch is one unit stop plus one unit start on the anchor); the
|
|
peers' `wg0.conf` digests are unchanged; the peers' `wg show wg0 latest-handshakes` advance
|
|
after the switch.
|
|
- `overlay show` lists `anchor` as the hub over the tunnel it took over, and both peers under
|
|
"peers of the tunnel … not nodes of the mesh", not yet enrolled.
|
|
- **T2 — a peer enrols and keeps its address.** On `peer-a`: `node add peer-a --adopted`,
|
|
token issued, `mesh-host enrol --token …` **with the broker reached over the tunnel** (the
|
|
broker address in the token is `10.10.0.1:<bus>`, which only the tunnel routes); then `overlay
|
|
place peer-a --site house` and a push. Assert: `node show peer-a` says a tunnel `wg0` was
|
|
found and `overlay show` puts `peer-a` at **10.10.0.2**; the carried-peers list now says
|
|
`enrolled as peer-a`; the anchor's `mesh0` still has exactly one entry for `peer-a`'s key;
|
|
`peer-a`'s `wg0` is down and `mesh0` up with the same key; `peer-a` still reaches
|
|
`10.10.0.1:8081` and `peer-b` still does too, uninterrupted.
|
|
- **T3 — a new machine gets a fresh address from the same range.** `fresh` enrols converged
|
|
(no tunnel), is placed, pushed. Assert its address is **10.10.0.4** (`.1` hub, `.2` and `.3`
|
|
the tunnel's), that it reaches `10.10.0.1` (the hub) and `10.10.0.2` (the enrolled peer) —
|
|
`ping -c1` over `mesh0` — and that `peer-b`, never enrolled, is still served.
|
|
- **T4 — nothing derived from the address is stale.** `plan anchor --json` and `plan peer-a
|
|
--json` (and the rendered `/etc/hosts` on each node) name `10.10.0.1` for the anchor and
|
|
`10.10.0.2` for `peer-a`, and no address in `10.42.0.0/16`; the broker address handed to a
|
|
module issued on `peer-a` is `anchor.internal:<bus>` resolving to `10.10.0.1`; the same after a
|
|
second `push` of every node, byte for byte.
|
|
- **N — the narrowed ADR 0100 check.** On `fresh`, a converged genesis dry-run with
|
|
`--overlay-range 10.10.0.0/24` while a *second* tunnel the bed raises there (`wg1` at
|
|
10.10.0.9/24, not adopted because the node is converged) is up, is refused naming `wg1` —
|
|
the non-overlap rule still applies where a tunnel is not adopted.
|
|
|
|
## What is not asserted here
|
|
|
|
- Taking a service over the tunnel (ADR 0100's bed does that).
|
|
- IPv6 tunnels: the parser reads them, the bed prepares only IPv4.
|