Files
jschoubben 4566c5c9aa Adopt the tunnel as a mesh fact, refuse a mismatched takeover, and rekey after enrolment
Review of the ADR 0105 build (hq ADR 0105). Four things it got wrong and one
path it lacked:

- A predecessor spoke's tunnel names one peer, the hub, routed the whole
  range; recording refused it and the whole enrolment failed. Range-routed
  peers are skipped now — only the hub's peers are ever carried.
- The range and the carried peers were conditions on the node being adopted,
  so converging the hub would have renumbered the mesh and dropped the peers
  still reaching it. They are facts of the tunnel record now, mode aside; the
  takeover alone is declared to an adopted node. Converging the hub is refused
  while a carried peer has not enrolled, naming it.
- A push composed a takeover for a hub whose address or endpoint disagreed
  with the tunnel, which would have the host stop the found interface and
  raise the mesh's where no peer listens. The graph refuses to compose it,
  naming both and the placement that fixes it.
- The host's account said taken or not; "found down and the mesh's not up"
  read as not taken. Three states now, and an account on every takeover.
- A hub that enrolled before this feature holds a key of its own, and
  re-enrolling would rotate every key the mesh sealed credentials to. A node
  now rekeys in a report, signed with its identity key over the key it
  leaves, the key it takes and the tunnel; the mesh verifies against the live
  key, refuses a stale or foreign proof, records key and tunnel, and moves a
  hub to the tunnel's address. `overlay show` names the path for a hub that
  found no tunnel.

Also: a carried IPv6 peer is routed /128, and identity.ForTest exists so the
link can be tested against a real identity store.
2026-09-24 00:02:07 +02:00

110 lines
7.7 KiB
Markdown

# Lab bed: the hub adopts the predecessor's tunnel (novox/hq ADR 0105)
A scenario and an integration-test skeleton for the mesh-lab repository, kept here because this
branch changes only the controller and the host. Move `adopt-the-tunnel.yml` to
`mesh-lab/scenarios/` and `adopt-the-tunnel.test.ts` to `mesh-lab/test/integration/` when the
feature lands; neither has been run. The skeleton follows `adoption.test.ts` and reuses its
harness. Documentation addresses throughout; the bed is node-agnostic.
## The bed, precisely
Three machines on one public segment, `hosting` (192.0.2.0/24), inbound allowed on all (the
anchor's firewall is the predecessor's, installed by the bed):
| machine | address | role |
|---|---|---|
| `anchor` | 192.0.2.10 | the machine in use: the predecessor's hub, then the mesh adopted on it |
| `peer-a` | 192.0.2.20 | a predecessor machine: reaches a service on the anchor through the tunnel; later **enrols and keeps its address** |
| `peer-b` | 192.0.2.30 | a second predecessor machine: reaches the same service; **never enrols** — the peer that must notice nothing throughout |
| `fresh` | 192.0.2.40 | a new machine: enrols later and **gets a fresh address from the same range** |
**Prepared the way the predecessor leaves a hub** (before genesis, by the bed, on `anchor`):
- `wireguard-tools` installed; a keypair made on each of `anchor`, `peer-a`, `peer-b`.
- `/etc/wireguard/wg0.conf` on the anchor: `[Interface]` `PrivateKey = <anchor's>`,
`ListenPort = 51900`, `Address = 10.10.0.1/24`; two `[Peer]` sections — `peer-a`'s public
key with `AllowedIPs = 10.10.0.2/32`, `peer-b`'s with `AllowedIPs = 10.10.0.3/32`. Raised with
`systemctl enable --now wg-quick@wg0`. **10.10.0.0/24 is deliberately not the mesh's default
range** (10.42.0.0/16), so a hub address in 10.10.0.0/24 can only have come from the tunnel.
- `wg0.conf` on each peer: its own key, `Address = 10.10.0.2/24` (resp. `.3/24`), one
`[Peer]` — the anchor's public key, `Endpoint = 192.0.2.10:51900`,
`AllowedIPs = 10.10.0.0/24`, `PersistentKeepalive = 25`. Raised the same way.
- A service on the anchor the peers reach **only over the tunnel**: a container publishing
`10.10.0.1:8081:80` (bound to the tunnel address, so a call from 192.0.2.20 to 10.10.0.1:8081
proves the tunnel carried it). Under a name no catalogue module uses — this bed is about the
tunnel, not about taking a service.
- The predecessor's firewall (`ufw`) allowing `51900/udp` and `22/tcp`, denying the rest — as ADR
0100's bed prepares it.
- A record of the anchor's `wg0` public key and of `sha256sum /etc/wireguard/wg0.conf`, taken
before genesis, for the assertions below.
**Genesis**, adopted, on the anchor: `mesh-bootstrap --adopted --node anchor --site hosting
--endpoint 192.0.2.10:51900 …` — no `--hub-port`, no `--overlay-range`, no `--tunnel`: the
installer finds `wg0` itself (one interface besides `mesh0`) and takes its port and range. The
bed asserts genesis **says** it found and took the tunnel.
**Review changes (2026-09-24).** A spoke's `wg0.conf` names one peer — the hub — routed the
whole range; the controller skips range-routed peers, so T2's enrolment carries no peer from the
spoke. A hub that enrolled *before* this feature (a generated key) takes the tunnel over without
re-enrolling: `mesh-host overlay take --tunnel wg0` on the machine rekeys the overlay key only and
sends a signed rekey; the bed adds **R0** for it below. A takeover is composed only for a hub
placed at the tunnel's address on the tunnel's port, and the host stops nothing until the declared
interface matches the found one and the key file holds the found key; a mesh interface that fails
to start gives the found unit back. The host's account has three states: `not-taken`, `taken`,
`down`.
## Assertions, in the record's order
- **R0 — a hub enrolled with its own key takes the tunnel over by rekeying.** Genesis is run
adopted *without* the tunnel being found (the bed stops `wg-quick@wg0` for the run, so the
installer sees no tunnel, then starts it again — the pre-feature shape). Then on the anchor:
`mesh-host overlay take --tunnel wg0`; `node show anchor` says "tunnel found wg0 …"; `overlay
place anchor --hub --endpoint 192.0.2.10:51900 --site hosting` (with `:51820` first, which must
be refused naming 51900); `plan anchor --json` names `Address = 10.10.0.1/32`, `ListenPort =
51900`, two `/32` peers, `takes-over` wg0 and nothing in 10.42.0.0/16; then `push anchor --wait
2m` and T1's assertions hold. `overlay take` run a second time is refused by the controller as
stale and changes nothing.
- **T1 — the tunnel changes hands and the peers notice nothing.** After genesis and the push
that raises the private network on the anchor:
- `wg show interfaces` on the anchor lists `mesh0` and not `wg0`;
`systemctl is-active wg-quick@wg0` is inactive and `is-enabled` disabled;
- `/etc/wireguard/wg0.conf` is on disk with the recorded digest (kept, never flushed), and
`node show anchor` names where its original was kept;
- `wg show mesh0 public-key` is the anchor's recorded `wg0` public key; `wg show mesh0
listen-port` is 51900; `ip -o addr show dev mesh0` carries `10.10.0.1`; `wg show mesh0 peers`
lists both peers' public keys with their `/32` allowed addresses;
- a loop on `peer-a` and `peer-b` calling `http://10.10.0.1:8081/` every second, started before
genesis, records **no window of failure longer than one WireGuard re-handshake** (measure and
assert an upper bound — the switch is one unit stop plus one unit start on the anchor); the
peers' `wg0.conf` digests are unchanged; the peers' `wg show wg0 latest-handshakes` advance
after the switch.
- `overlay show` lists `anchor` as the hub over the tunnel it took over, and both peers under
"peers of the tunnel … not nodes of the mesh", not yet enrolled.
- **T2 — a peer enrols and keeps its address.** On `peer-a`: `node add peer-a --adopted`,
token issued, `mesh-host enrol --token …` **with the broker reached over the tunnel** (the
broker address in the token is `10.10.0.1:<bus>`, which only the tunnel routes); then `overlay
place peer-a --site house` and a push. Assert: `node show peer-a` says a tunnel `wg0` was
found and `overlay show` puts `peer-a` at **10.10.0.2**; the carried-peers list now says
`enrolled as peer-a`; the anchor's `mesh0` still has exactly one entry for `peer-a`'s key;
`peer-a`'s `wg0` is down and `mesh0` up with the same key; `peer-a` still reaches
`10.10.0.1:8081` and `peer-b` still does too, uninterrupted.
- **T3 — a new machine gets a fresh address from the same range.** `fresh` enrols converged
(no tunnel), is placed, pushed. Assert its address is **10.10.0.4** (`.1` hub, `.2` and `.3`
the tunnel's), that it reaches `10.10.0.1` (the hub) and `10.10.0.2` (the enrolled peer) —
`ping -c1` over `mesh0` — and that `peer-b`, never enrolled, is still served.
- **T4 — nothing derived from the address is stale.** `plan anchor --json` and `plan peer-a
--json` (and the rendered `/etc/hosts` on each node) name `10.10.0.1` for the anchor and
`10.10.0.2` for `peer-a`, and no address in `10.42.0.0/16`; the broker address handed to a
module issued on `peer-a` is `anchor.internal:<bus>` resolving to `10.10.0.1`; the same after a
second `push` of every node, byte for byte.
- **N — the narrowed ADR 0100 check.** On `fresh`, a converged genesis dry-run with
`--overlay-range 10.10.0.0/24` while a *second* tunnel the bed raises there (`wg1` at
10.10.0.9/24, not adopted because the node is converged) is up, is refused naming `wg1` —
the non-overlap rule still applies where a tunnel is not adopted.
## What is not asserted here
- Taking a service over the tunnel (ADR 0100's bed does that).
- IPv6 tunnels: the parser reads them, the bed prepares only IPv4.