--- status: accepted date: 2026-08-22 deciders: jochen reconstructed: false --- # 16. A lab node is a virtual machine running the real install ## Context [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) makes a local mesh a prerequisite rather than a convenience: *"everything that manifests between nodes is discoverable only in production, which is where every fault of 2026-08-22 was found."* Two things were measured while establishing what exists ([`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md)): - **There is no local mesh.** The dev tooling starts providers through the host's own init system and reads credentials from host paths (`modules/hal/developer/tools/dev-env.ts:146-178`). It borrows the machine because there is nowhere else to put a mesh. - **The one containerised node in the repository has been unable to build since 2026-06-04**, when the npm workspace it depends on was removed. Nothing runs it, so nothing reported it. So the question is not how to improve a local mesh. It is what a node *is* when it is not a physical machine. Every subsequent question — how faithful is faithful enough, what may be mocked, which failures remain reachable — follows from that one answer. The hardware available is not a constraint: 125 GB of memory with 71 free, 24 threads, and hardware virtualisation present. ## Considered Options 1. **An application container.** Rejected. **A node's job is to run containers**, so modelling a node as one inverts the thing being modelled: module service stacks then require nested containers through a privileged daemon, or a shared socket that makes isolation between nodes cosmetic. Init is not PID 1, so units and timers need workarounds. Cheapest to start and the least like a node. 2. **A system container.** Rejected, after first being recommended. It is genuinely good — real init, properly nested containers, roughly a second to boot, cheap snapshots — and it is the only option that makes a twenty-node run affordable. It was rejected because **the scale requirement that justified it was invented rather than required**: the stated goal is to run the real mesh, which is four nodes, on one computer. And a system container still forces the question a virtual machine dissolves — *how faithful must a node be?* — which then has to be answered again for every capability under test. 3. **`systemd-nspawn`.** Rejected. Already present, so nothing to install, but too primitive: no storage pools, no snapshot management, no network management, no virtual machines. Snapshots are what make the loop fast, so the saving is not worth what it costs. 4. **A virtual machine.** **Adopted.** A bare Arch Linux machine that the real install script turns into a node. ## Decision **A node in the mesh development lab is a virtual machine.** It boots a stock Linux image, runs the real install, and becomes a node. It is not a model of a node, so no question arises about how good the model is. The environment is called **the lab**. Three things follow directly and are decided here: ### The lab is driven by `incus` Chosen for what it manages, not for what it is: virtual machines, their snapshots, and the bridges between them, through one interface. It also manages system containers, so if a run ever genuinely needs twenty nodes, that is a change of instance type rather than a rewrite. Declared in `modules/hal/developer/module.yml`, so it installs the way every other package does. ### The simulated public segment uses TEST-NET-3 `203.0.113.0/24`, reserved by RFC 5737, never routable. This is not cosmetic. WireGuard decides per pair whether to write an `Endpoint` by testing the peer's underlay address against an RFC1918 regex (`modules/wireguard/hooks/index.ts:225-240`). A simulated public segment addressed from private space makes the hub test as unreachable, so no spoke writes an endpoint for it, nothing can initiate, **and the mesh silently never forms** — appearing as a WireGuard fault rather than an addressing mistake. The production LAN subnet and the entire overlay address plan are reproduced unchanged. ### The lab issues its own certificates Public names are certified by an ACME server inside the lab; `.internal` names keep the mesh CA. **The lab keeps production's two-authority split rather than collapsing it**, because a single-authority lab would hide any fault living in that split. This also makes the lab's port forward load-bearing: an HTTP-01 challenge must reach a published-but-NATed node on port 80, so a broken forward becomes a reproducible certificate failure rather than a mystery. ## Consequences **The fidelity question disappears, and with it a class of argument.** There is no "how real is this node" to litigate per capability, because the node is real. What remains not-real is a short, enumerable list: the model provider, the public internet, and the public certificate authority. **The install becomes the thing under test.** A container-shaped lab would have had to skip the bootstrap entirely. Here it runs, so it is exercised on every fresh lab. **Reproducing the network is mostly a data problem.** The bootstrap performs no network configuration at all; WireGuard, DNS, routing and internal TLS are generated by module hooks from mesh-DB rows. The lab therefore exercises the same code production runs rather than a reimplementation ([`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md)). **Scale runs get expensive, and this is the real cost.** Four virtual machines are comfortable; twenty are not, on a workstation. Faults that only appear at scale — a fan-out reaching most consumers rather than all, a cascade that stalls with many modules — stay hard to reproduce. The mitigation is that the same tooling runs system containers, so a scale run remains possible at lower fidelity if one is ever genuinely needed. **Boot is slower, and it does not matter.** Ten to twenty seconds against roughly one. A run includes a full delivery — build, publish, install, migrate — measured in minutes, so boot time is noise. **One change is required before the lab can issue certificates.** The reverse proxy sets no `caServer`, so it defaults to the public authority's *production* endpoint (`modules/traefik/docker-compose.yml:17-19`). It must become configurable, defaulting to production so real nodes are unaffected. Worth noting on its own: aiming at production rather than staging means every certificate experiment on a real node consumes issuance quota. ## References - [`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md) — what exists, and the four host couplings that only obstruct a container-shaped node - [`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md) — the topology being reproduced and the endpoint constraint - [`02-DESIGN/01-end-to-end-testing.md`](../02-DESIGN/01-to-be/01-end-to-end-testing.md) — what the lab is for - `modules/wireguard/hooks/index.ts:206-240` — the endpoint rule, and the incident comments recording what it cost to get right - RFC 5737 — reserved documentation address blocks