Files
hq/02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
T
jschoubben c0b35652d0 The numbering is the flow: decisions are 02, design is 03
papa-hq reads 01 research -> 03 decision -> 02 design. The order is a
scar, not a choice: 02-DESIGN existed from its initial commit, and when
adr/ was finally promoted on 2026-07-13 it took the next free number
rather than its place in the sequence. By then design was too settled to
renumber.

hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and
02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks
the process in the order it happens: research produces a decision, the
decision authorises a design.

00-GENESIS becomes 00-META, matching papa's rename from the same
restructure.

Every path reference rewritten across documents, frontmatter, playbooks
and skills. All links resolve; all 58 frontmatter blocks parse and their
path fields still point at files that exist.
2026-08-23 18:05:11 +02:00

139 lines
7.1 KiB
Markdown

---
status: accepted
date: 2026-08-22
deciders: jochen
reconstructed: false
---
# 16. A lab node is a virtual machine running the real install
## Context
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) makes a local mesh a prerequisite
rather than a convenience: *"everything that manifests between nodes is discoverable only in
production, which is where every fault of 2026-08-22 was found."*
Two things were measured while establishing what exists
([`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md)):
- **There is no local mesh.** The dev tooling starts providers through the host's own init
system and reads credentials from host paths (`modules/hal/developer/tools/dev-env.ts:146-178`).
It borrows the machine because there is nowhere else to put a mesh.
- **The one containerised node in the repository has been unable to build since 2026-06-04**,
when the npm workspace it depends on was removed. Nothing runs it, so nothing reported it.
So the question is not how to improve a local mesh. It is what a node *is* when it is not a
physical machine. Every subsequent question — how faithful is faithful enough, what may be
mocked, which failures remain reachable — follows from that one answer.
The hardware available is not a constraint: 125 GB of memory with 71 free, 24 threads, and
hardware virtualisation present.
## Considered Options
1. **An application container.** Rejected. **A node's job is to run containers**, so modelling
a node as one inverts the thing being modelled: module service stacks then require nested
containers through a privileged daemon, or a shared socket that makes isolation between
nodes cosmetic. Init is not PID 1, so units and timers need workarounds. Cheapest to start
and the least like a node.
2. **A system container.** Rejected, after first being recommended. It is genuinely good —
real init, properly nested containers, roughly a second to boot, cheap snapshots — and it
is the only option that makes a twenty-node run affordable. It was rejected because **the
scale requirement that justified it was invented rather than required**: the stated goal is
to run the real mesh, which is four nodes, on one computer. And a system container still
forces the question a virtual machine dissolves — *how faithful must a node be?* — which
then has to be answered again for every capability under test.
3. **`systemd-nspawn`.** Rejected. Already present, so nothing to install, but too primitive:
no storage pools, no snapshot management, no network management, no virtual machines.
Snapshots are what make the loop fast, so the saving is not worth what it costs.
4. **A virtual machine.** **Adopted.** A bare Arch Linux machine that the real install script
turns into a node.
## Decision
**A node in the mesh development lab is a virtual machine.** It boots a stock Linux image,
runs the real install, and becomes a node. It is not a model of a node, so no question arises
about how good the model is.
The environment is called **the lab**.
Three things follow directly and are decided here:
### The lab is driven by `incus`
Chosen for what it manages, not for what it is: virtual machines, their snapshots, and the
bridges between them, through one interface. It also manages system containers, so if a run
ever genuinely needs twenty nodes, that is a change of instance type rather than a rewrite.
Declared in `modules/hal/developer/module.yml`, so it installs the way every other package
does.
### The simulated public segment uses TEST-NET-3
`203.0.113.0/24`, reserved by RFC 5737, never routable.
This is not cosmetic. WireGuard decides per pair whether to write an `Endpoint` by testing the
peer's underlay address against an RFC1918 regex
(`modules/wireguard/hooks/index.ts:225-240`). A simulated public segment addressed from
private space makes the hub test as unreachable, so no spoke writes an endpoint for it,
nothing can initiate, **and the mesh silently never forms** — appearing as a WireGuard fault
rather than an addressing mistake.
The production LAN subnet and the entire overlay address plan are reproduced unchanged.
### The lab issues its own certificates
Public names are certified by an ACME server inside the lab; `.internal` names keep the mesh
CA. **The lab keeps production's two-authority split rather than collapsing it**, because a
single-authority lab would hide any fault living in that split.
This also makes the lab's port forward load-bearing: an HTTP-01 challenge must reach a
published-but-NATed node on port 80, so a broken forward becomes a reproducible certificate
failure rather than a mystery.
## Consequences
**The fidelity question disappears, and with it a class of argument.** There is no "how real
is this node" to litigate per capability, because the node is real. What remains not-real is a
short, enumerable list: the model provider, the public internet, and the public certificate
authority.
**The install becomes the thing under test.** A container-shaped lab would have had to skip
the bootstrap entirely. Here it runs, so it is exercised on every fresh lab.
**Reproducing the network is mostly a data problem.** The bootstrap performs no network
configuration at all; WireGuard, DNS, routing and internal TLS are generated by module hooks
from mesh-DB rows. The lab therefore exercises the same code production runs rather than a
reimplementation ([`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md)).
**Scale runs get expensive, and this is the real cost.** Four virtual machines are
comfortable; twenty are not, on a workstation. Faults that only appear at scale — a fan-out
reaching most consumers rather than all, a cascade that stalls with many modules — stay hard
to reproduce. The mitigation is that the same tooling runs system containers, so a scale run
remains possible at lower fidelity if one is ever genuinely needed.
**Boot is slower, and it does not matter.** Ten to twenty seconds against roughly one. A run
includes a full delivery — build, publish, install, migrate — measured in minutes, so boot
time is noise.
**One change is required before the lab can issue certificates.** The reverse proxy sets no
`caServer`, so it defaults to the public authority's *production* endpoint
(`modules/traefik/docker-compose.yml:17-19`). It must become configurable, defaulting to
production so real nodes are unaffected. Worth noting on its own: aiming at production rather
than staging means every certificate experiment on a real node consumes issuance quota.
## References
- [`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md) — what exists, and
the four host couplings that only obstruct a container-shaped node
- [`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md) — the topology
being reproduced and the endpoint constraint
- [`03-DESIGN/01-end-to-end-testing.md`](../03-DESIGN/01-to-be/01-end-to-end-testing.md) — what the lab
is for
- `modules/wireguard/hooks/index.ts:206-240` — the endpoint rule, and the incident comments
recording what it cost to get right
- RFC 5737 — reserved documentation address blocks