Files
hq/adr/0016-a-lab-node-is-a-virtual-machine.md
T
jschoubben 702efca6bb Base layer: the mesh as it is, under the mesh as it should be
HQ held only the to-be. Every reader had to already know the system the
decisions were about, and an as-is claim had nowhere to live except inside
an intention.

Adds 02-DESIGN/00-as-is — eleven documents written from the implementation
and the operational record, not from intent, including the parts nobody
would choose again. The two existing designs move under 01-to-be. Layers
are declared in frontmatter and never mix: a design that ships does not
move, its as-is counterpart is written, and both stand.

Back-fills adr/0001-0014 for decisions taken in implementation and never
recorded — the broker, the module abstraction, the mesh database, managed
files, provisioning, migrations, the workspace removal, failing loudly,
the constitution, application placement, linking, the employee model, the
artifact, the three silos. Each marked reconstructed, dated from the
history, and citing the evidence it was recovered from. The two existing
records renumber to 0015 and 0016 so the ledger runs oldest first;
0017 extends 0015 to modules outside the core, principle only — the
domain list is deliberately not invented here.

how-we-build.md becomes the source of the mesh constitution, with a sync
playbook, so the enforced copy stops being the only one that is true.

Process becomes explicit: five playbooks, eight thin skills that defer to
them, a repository map, and AGENTS.md with CLAUDE.md as its include.

The five Observations become 04-ISSUES 001-005 where they can be owned and
closed. 006 is new and uncomfortable: HQ is not indexed into the knowledge
base. That claim is what decision 27 rests on, it was never checked, and
the README now says so instead of repeating it.

Also corrects the ADR index into something generated, the "02-DESIGN is
empty" claim, the VISION.md pointer that did not survive the repo split,
and a note asserting the symlink rule was contradicted — it was a
misreading; the rule forbids hand-made links, the installer links by design.
2026-08-23 03:08:26 +02:00

7.1 KiB

status, date, deciders, reconstructed
status date deciders reconstructed
accepted 2026-08-22 jochen false

16. A lab node is a virtual machine running the real install

Context

ADR 0015 makes a local mesh a prerequisite rather than a convenience: "everything that manifests between nodes is discoverable only in production, which is where every fault of 2026-08-22 was found."

Two things were measured while establishing what exists (01-RESEARCH/002-local-mesh):

  • There is no local mesh. The dev tooling starts providers through the host's own init system and reads credentials from host paths (modules/hal/developer/tools/dev-env.ts:146-178). It borrows the machine because there is nowhere else to put a mesh.
  • The one containerised node in the repository has been unable to build since 2026-06-04, when the npm workspace it depends on was removed. Nothing runs it, so nothing reported it.

So the question is not how to improve a local mesh. It is what a node is when it is not a physical machine. Every subsequent question — how faithful is faithful enough, what may be mocked, which failures remain reachable — follows from that one answer.

The hardware available is not a constraint: 125 GB of memory with 71 free, 24 threads, and hardware virtualisation present.

Considered Options

  1. An application container. Rejected. A node's job is to run containers, so modelling a node as one inverts the thing being modelled: module service stacks then require nested containers through a privileged daemon, or a shared socket that makes isolation between nodes cosmetic. Init is not PID 1, so units and timers need workarounds. Cheapest to start and the least like a node.

  2. A system container. Rejected, after first being recommended. It is genuinely good — real init, properly nested containers, roughly a second to boot, cheap snapshots — and it is the only option that makes a twenty-node run affordable. It was rejected because the scale requirement that justified it was invented rather than required: the stated goal is to run the real mesh, which is four nodes, on one computer. And a system container still forces the question a virtual machine dissolves — how faithful must a node be? — which then has to be answered again for every capability under test.

  3. systemd-nspawn. Rejected. Already present, so nothing to install, but too primitive: no storage pools, no snapshot management, no network management, no virtual machines. Snapshots are what make the loop fast, so the saving is not worth what it costs.

  4. A virtual machine. Adopted. A bare Arch Linux machine that the real install script turns into a node.

Decision

A node in the mesh development lab is a virtual machine. It boots a stock Linux image, runs the real install, and becomes a node. It is not a model of a node, so no question arises about how good the model is.

The environment is called the lab.

Three things follow directly and are decided here:

The lab is driven by incus

Chosen for what it manages, not for what it is: virtual machines, their snapshots, and the bridges between them, through one interface. It also manages system containers, so if a run ever genuinely needs twenty nodes, that is a change of instance type rather than a rewrite.

Declared in modules/hal/developer/module.yml, so it installs the way every other package does.

The simulated public segment uses TEST-NET-3

203.0.113.0/24, reserved by RFC 5737, never routable.

This is not cosmetic. WireGuard decides per pair whether to write an Endpoint by testing the peer's underlay address against an RFC1918 regex (modules/wireguard/hooks/index.ts:225-240). A simulated public segment addressed from private space makes the hub test as unreachable, so no spoke writes an endpoint for it, nothing can initiate, and the mesh silently never forms — appearing as a WireGuard fault rather than an addressing mistake.

The production LAN subnet and the entire overlay address plan are reproduced unchanged.

The lab issues its own certificates

Public names are certified by an ACME server inside the lab; .internal names keep the mesh CA. The lab keeps production's two-authority split rather than collapsing it, because a single-authority lab would hide any fault living in that split.

This also makes the lab's port forward load-bearing: an HTTP-01 challenge must reach a published-but-NATed node on port 80, so a broken forward becomes a reproducible certificate failure rather than a mystery.

Consequences

The fidelity question disappears, and with it a class of argument. There is no "how real is this node" to litigate per capability, because the node is real. What remains not-real is a short, enumerable list: the model provider, the public internet, and the public certificate authority.

The install becomes the thing under test. A container-shaped lab would have had to skip the bootstrap entirely. Here it runs, so it is exercised on every fresh lab.

Reproducing the network is mostly a data problem. The bootstrap performs no network configuration at all; WireGuard, DNS, routing and internal TLS are generated by module hooks from mesh-DB rows. The lab therefore exercises the same code production runs rather than a reimplementation (01-RESEARCH/004-lab-network).

Scale runs get expensive, and this is the real cost. Four virtual machines are comfortable; twenty are not, on a workstation. Faults that only appear at scale — a fan-out reaching most consumers rather than all, a cascade that stalls with many modules — stay hard to reproduce. The mitigation is that the same tooling runs system containers, so a scale run remains possible at lower fidelity if one is ever genuinely needed.

Boot is slower, and it does not matter. Ten to twenty seconds against roughly one. A run includes a full delivery — build, publish, install, migrate — measured in minutes, so boot time is noise.

One change is required before the lab can issue certificates. The reverse proxy sets no caServer, so it defaults to the public authority's production endpoint (modules/traefik/docker-compose.yml:17-19). It must become configurable, defaulting to production so real nodes are unaffected. Worth noting on its own: aiming at production rather than staging means every certificate experiment on a real node consumes issuance quota.

References

  • 01-RESEARCH/002-local-mesh — what exists, and the four host couplings that only obstruct a container-shaped node
  • 01-RESEARCH/004-lab-network — the topology being reproduced and the endpoint constraint
  • 02-DESIGN/01-end-to-end-testing.md — what the lab is for
  • modules/wireguard/hooks/index.ts:206-240 — the endpoint rule, and the incident comments recording what it cost to get right
  • RFC 5737 — reserved documentation address blocks