Files
hq/02-DECISIONS/0016-the-lab.md
T
jschoubben 53b94c51bb The pointers back from what yesterday's records changed, which I missed twice
Three records were left describing a mechanism a new record had moved,
and a reader arrives at them by following a citation: 0066 still said a
routed name is written into every container after 0148 replaced that with
resolution; 0016 still read as though the lab were the test bed after
0149; and issues 109 and 135 said nothing about 0148 ending the copying
that 135's own fix made comparable. Each was a citation leading to the
wrong answer in a record that was not wrong about anything it decided.

This is the second time in one session. The playbook rule I added last
round did not stop it, so the convention is now written where the record
conventions live, with the shape to use and three worked examples — and
with the honest note that it is NOT machine-checked and cannot be from
`extends:` alone: 102 records extend another, 87 have no back-reference,
and that is correct, because extending usually means building on a
context. Making it mechanical means a record declaring the relationship
in frontmatter, which is a schema change and is not mine to decide.

Also: designs 18 and 20 claimed `updated:` dates from before I edited
them, and 117's `fixed-by` gained the commit beside the record.
2026-09-30 00:47:19 +02:00

93 lines
4.6 KiB
Markdown

---
topic: building it
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 16. The lab
*Consolidated 2026-08-28 from five records. The lab is one design and was split across five
decisions taken over three days; the reasoning is kept, the fragmentation is not.*
The environment a change is run against before it reaches real machines.
> **Still the lab, no longer the test bed — 2026-09-30, by [ADR 0149](0149-the-live-mesh-is-the-test-bed.md).**
> Everything here stands. What changed is what the lab is *for*: a change is verified against the mesh
> that is running, because the faults that cost the most are faults of a mesh that already exists —
> bound consumers, containers made against an older roster, an adopted machine — and a bed is by
> construction a mesh that does not. Raising a mesh from bare is now the lab's whole job, which is the
> one thing the live mesh cannot be asked to do. 0149 also supersedes
> [ADR 0068](0068-the-lab-takes-requests.md), which extended this one and was never built.
## A node in the lab is a virtual machine
It boots a stock Linux image, runs the real install, and becomes a node. **It is not a model of
a node**, so no question arises about how good the model is — which is the whole reason for
paying the cost of virtual machines rather than containers.
The lab is driven by **incus**, and a scenario is raised from a declaration.
## A router is scenery, and is therefore a container
**Nothing under test runs on a router.** It is not a participant, holds no identity, has nothing
installed on it by the mesh, and no assertion is ever made about its internals. It exists so that
packets between machines behave the way they behave in the world.
The fidelity argument that makes a node a virtual machine does not reach it: what a router *is*
does not matter, only what it *does to traffic*. So a router is a system container, and the lab
is cheaper for it.
## A scenario declares the underlay, and only the underlay
**What a hosting provider and a home router would have provided**, before any of our software
touched the machine:
- which segments exist, and their address ranges
- which machine sits on which segment, at which address
- what NAT sits between them, and which ports are forwarded through it
- which machines are detached, and may be attached or detached during a run
**A scenario declares nothing about the overlay** — no overlay addresses, no hub, no peering, no
names, no certificates. Those are the mesh's job, and a scenario that supplied them would be
testing itself.
> A scenario provides what a hosting provider and a home router would provide, and nothing our
> software is responsible for.
## A scenario is a closed address space
Every segment materialises as its own isolated link belonging to one scenario instance. **Two
scenarios raised from the same declaration hold the same addresses and never meet**, because
nothing joins their links. The declaration therefore keeps its literal addresses and they mean
exactly what they say.
**The consequence that constrains everything else: the lab never reaches into a scenario over
IP.** It talks to a machine through the virtualisation layer's own channel — the way one would
use a console rather than the network. That is what makes two identical scenarios able to run at
once, and it is why placing anything inside a machine is a hypervisor operation rather than a
network one.
## Two scenario classes, and the first has no pipeline
| | **bootstrap** | **full** |
|---|---|---|
| contains | machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules |
| verdict from | what the host reports about the state it reconciled | a delivery result ending in verification |
| exercises | tiers 0 and 1 | tiers 2 and above, and modules |
**The bootstrap class comes first**, because it is what develops the node host, and because a
full scenario needs tiers that do not exist yet. A lab that could only raise the larger class
would be a lab nobody could use until everything else was built.
## Consequences
- **The lab tests the real code path**, not a reimplementation of it. The network a scenario
produces is generated by the same code production runs.
- **Isolation is what makes it usable in parallel**, and it costs the ability to reach in over
IP. Everything the lab puts inside a machine — a binary, an image, a file — goes through the
hypervisor.
- **A sealed scenario cannot fetch anything**, which is a real limit rather than an inconvenience:
it is why images have to be placed and why a container runtime has to be in the base image.