Files
hq/02-DECISIONS/0016-the-lab.md
T
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00

3.9 KiB

status, date, deciders, reconstructed
status date deciders reconstructed
accepted 2026-08-28 jochen false

16. The lab

Consolidated 2026-08-28 from five records. The lab is one design and was split across five decisions taken over three days; the reasoning is kept, the fragmentation is not.

The environment a change is run against before it reaches real machines.

A node in the lab is a virtual machine

It boots a stock Linux image, runs the real install, and becomes a node. It is not a model of a node, so no question arises about how good the model is — which is the whole reason for paying the cost of virtual machines rather than containers.

The lab is driven by incus, and a scenario is raised from a declaration.

A router is scenery, and is therefore a container

Nothing under test runs on a router. It is not a participant, holds no identity, has nothing installed on it by the mesh, and no assertion is ever made about its internals. It exists so that packets between machines behave the way they behave in the world.

The fidelity argument that makes a node a virtual machine does not reach it: what a router is does not matter, only what it does to traffic. So a router is a system container, and the lab is cheaper for it.

A scenario declares the underlay, and only the underlay

What a hosting provider and a home router would have provided, before any of our software touched the machine:

  • which segments exist, and their address ranges
  • which machine sits on which segment, at which address
  • what NAT sits between them, and which ports are forwarded through it
  • which machines are detached, and may be attached or detached during a run

A scenario declares nothing about the overlay — no overlay addresses, no hub, no peering, no names, no certificates. Those are the mesh's job, and a scenario that supplied them would be testing itself.

A scenario provides what a hosting provider and a home router would provide, and nothing our software is responsible for.

A scenario is a closed address space

Every segment materialises as its own isolated link belonging to one scenario instance. Two scenarios raised from the same declaration hold the same addresses and never meet, because nothing joins their links. The declaration therefore keeps its literal addresses and they mean exactly what they say.

The consequence that constrains everything else: the lab never reaches into a scenario over IP. It talks to a machine through the virtualisation layer's own channel — the way one would use a console rather than the network. That is what makes two identical scenarios able to run at once, and it is why placing anything inside a machine is a hypervisor operation rather than a network one.

Two scenario classes, and the first has no pipeline

bootstrap full
contains machines, the host binary, a pinned substrate bundle a complete mesh: forge, coordinator, delivery, modules
verdict from what the host reports about the state it reconciled a delivery result ending in verification
exercises tiers 0 and 1 tiers 2 and above, and modules

The bootstrap class comes first, because it is what develops the node host, and because a full scenario needs tiers that do not exist yet. A lab that could only raise the larger class would be a lab nobody could use until everything else was built.

Consequences

  • The lab tests the real code path, not a reimplementation of it. The network a scenario produces is generated by the same code production runs.
  • Isolation is what makes it usable in parallel, and it costs the ability to reach in over IP. Everything the lab puts inside a machine — a binary, an image, a file — goes through the hypervisor.
  • A sealed scenario cannot fetch anything, which is a real limit rather than an inconvenience: it is why images have to be placed and why a container runtime has to be in the base image.