Files
hq/adr/0002-a-lab-node-is-a-virtual-machine.md
T
jschoubben cf9357e8e9 HQ — the mesh's own documentation
What the mesh is, what it is becoming, and why. Implementation lives in the
code repositories; the reasoning lives here.

  00-GENESIS   mission, engineering context, effect, and the rules that hold
  01-RESEARCH  investigations, before they harden into design
  02-DESIGN    the authoritative specification
  adr          numbered decisions — what was chosen, and what was rejected
  DECISIONS.md the ledger: every decision, in the order it was taken

Written for a reader who is not its author and has no access to the mesh it
describes. Addresses use the documentation ranges of RFC 5737 and RFC 1918;
nodes are named by role.

Single initial commit by intent. The prior history came from a private
repository and carried operational detail — a routable address identified as a
VPN hub, real domain names, a hosting provider — which sanitising a tip commit
would not have removed from the log.
2026-08-22 22:01:32 +02:00

7.1 KiB

2. A lab node is a virtual machine running the real install

  • Status: Accepted
  • Date: 2026-08-22
  • Deciders: jochen

Context

ADR 0001 makes a local mesh a prerequisite rather than a convenience: "everything that manifests between nodes is discoverable only in production, which is where every fault of 2026-08-22 was found."

Two things were measured while establishing what exists (01-RESEARCH/002-local-mesh):

  • There is no local mesh. The dev tooling starts providers through the host's own init system and reads credentials from host paths (modules/hal/developer/tools/dev-env.ts:146-178). It borrows the machine because there is nowhere else to put a mesh.
  • The one containerised node in the repository has been unable to build since 2026-06-04, when the npm workspace it depends on was removed. Nothing runs it, so nothing reported it.

So the question is not how to improve a local mesh. It is what a node is when it is not a physical machine. Every subsequent question — how faithful is faithful enough, what may be mocked, which failures remain reachable — follows from that one answer.

The hardware available is not a constraint: 125 GB of memory with 71 free, 24 threads, and hardware virtualisation present.

Considered Options

  1. An application container. Rejected. A node's job is to run containers, so modelling a node as one inverts the thing being modelled: module service stacks then require nested containers through a privileged daemon, or a shared socket that makes isolation between nodes cosmetic. Init is not PID 1, so units and timers need workarounds. Cheapest to start and the least like a node.

  2. A system container. Rejected, after first being recommended. It is genuinely good — real init, properly nested containers, roughly a second to boot, cheap snapshots — and it is the only option that makes a twenty-node run affordable. It was rejected because the scale requirement that justified it was invented rather than required: the stated goal is to run the real mesh, which is four nodes, on one computer. And a system container still forces the question a virtual machine dissolves — how faithful must a node be? — which then has to be answered again for every capability under test.

  3. systemd-nspawn. Rejected. Already present, so nothing to install, but too primitive: no storage pools, no snapshot management, no network management, no virtual machines. Snapshots are what make the loop fast, so the saving is not worth what it costs.

  4. A virtual machine. Adopted. A bare Arch Linux machine that the real install script turns into a node.

Decision

A node in the mesh development lab is a virtual machine. It boots a stock Linux image, runs the real install, and becomes a node. It is not a model of a node, so no question arises about how good the model is.

The environment is called the lab.

Three things follow directly and are decided here:

The lab is driven by incus

Chosen for what it manages, not for what it is: virtual machines, their snapshots, and the bridges between them, through one interface. It also manages system containers, so if a run ever genuinely needs twenty nodes, that is a change of instance type rather than a rewrite.

Declared in modules/hal/developer/module.yml, so it installs the way every other package does.

The simulated public segment uses TEST-NET-3

203.0.113.0/24, reserved by RFC 5737, never routable.

This is not cosmetic. WireGuard decides per pair whether to write an Endpoint by testing the peer's underlay address against an RFC1918 regex (modules/wireguard/hooks/index.ts:225-240). A simulated public segment addressed from private space makes the hub test as unreachable, so no spoke writes an endpoint for it, nothing can initiate, and the mesh silently never forms — appearing as a WireGuard fault rather than an addressing mistake.

The production LAN subnet and the entire overlay address plan are reproduced unchanged.

The lab issues its own certificates

Public names are certified by an ACME server inside the lab; .internal names keep the mesh CA. The lab keeps production's two-authority split rather than collapsing it, because a single-authority lab would hide any fault living in that split.

This also makes the lab's port forward load-bearing: an HTTP-01 challenge must reach a published-but-NATed node on port 80, so a broken forward becomes a reproducible certificate failure rather than a mystery.

Consequences

The fidelity question disappears, and with it a class of argument. There is no "how real is this node" to litigate per capability, because the node is real. What remains not-real is a short, enumerable list: the model provider, the public internet, and the public certificate authority.

The install becomes the thing under test. A container-shaped lab would have had to skip the bootstrap entirely. Here it runs, so it is exercised on every fresh lab.

Reproducing the network is mostly a data problem. The bootstrap performs no network configuration at all; WireGuard, DNS, routing and internal TLS are generated by module hooks from mesh-DB rows. The lab therefore exercises the same code production runs rather than a reimplementation (01-RESEARCH/004-lab-network).

Scale runs get expensive, and this is the real cost. Four virtual machines are comfortable; twenty are not, on a workstation. Faults that only appear at scale — a fan-out reaching most consumers rather than all, a cascade that stalls with many modules — stay hard to reproduce. The mitigation is that the same tooling runs system containers, so a scale run remains possible at lower fidelity if one is ever genuinely needed.

Boot is slower, and it does not matter. Ten to twenty seconds against roughly one. A run includes a full delivery — build, publish, install, migrate — measured in minutes, so boot time is noise.

One change is required before the lab can issue certificates. The reverse proxy sets no caServer, so it defaults to the public authority's production endpoint (modules/traefik/docker-compose.yml:17-19). It must become configurable, defaulting to production so real nodes are unaffected. Worth noting on its own: aiming at production rather than staging means every certificate experiment on a real node consumes issuance quota.

References

  • 01-RESEARCH/002-local-mesh — what exists, and the four host couplings that only obstruct a container-shaped node
  • 01-RESEARCH/004-lab-network — the topology being reproduced and the endpoint constraint
  • 02-DESIGN/01-end-to-end-testing.md — what the lab is for
  • modules/wireguard/hooks/index.ts:206-240 — the endpoint rule, and the incident comments recording what it cost to get right
  • RFC 5737 — reserved documentation address blocks