Files
hq/03-DESIGN/00-as-is/11-the-lab.md
T
jschoubben 4bf7a35568 Close the record on the lab
Playbook 02 and 04 were followed for the substance — decisions before design,
design before build — and skipped for the bookkeeping. This closes that.

004 graduates. Its one open item was "not yet stood up"; the lab is stood up,
and the substitution the effort turned on is now enforced by the validator
before anything is raised rather than left as a thing to remember. Its
certificate conclusion has a home in 01-end-to-end-testing and is designed but
not built — implementation is a third axis, and an effort graduates on its
conclusions.

One item leaves 004 without a home and is recorded rather than lost: the reverse
proxy does not set caServer, so it defaults to the production endpoint.

The two lab designs read `designed` while running in production of a sort, so
they become `in-progress`.

And the lab gets an as-is document, which it did not have. It records what runs
including the parts nobody would choose again: that `place:` is refused and the
lab therefore raises EMPTY MACHINES, that the drawing shipped with no design
document behind it, that a router is tagged as a machine for a reason found by a
bug, and that the integration suite raises two of five scenarios while both
faults found so far lived in the three it does not.

006 stays active, deliberately. Two of its open questions ARE the tier 0 design
— whether absorbing six concerns makes the host too large, and whether an
unprivileged node earns a place in the inventory. Playbook 04 is explicit that
an open question is a reason to research, not to build around.
2026-08-25 01:55:52 +02:00

7.2 KiB
Raw Blame History

layer, status, code, updated, decisions
layer status code updated decisions
as-is implemented
mesh-lab
2026-08-25
02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
02-DECISIONS/0031-the-lab-provides-the-underlay.md
02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md
02-DECISIONS/0033-a-router-is-scenery-not-a-node.md

The lab, as it stands

The first piece of the new shape that exists. It is the only repository outside the monorepo so far, and unlike everything else planned it ships to nobody: it runs on a workstation, raises virtual machines, and throws them away.

Written from the implementation. Where intent and implementation disagree, the implementation is what is recorded here and the disagreement is stated.

What it does

A scenario is a YAML file declaring an underlay — segments, the gateways between them, and machines placed on them. raise materialises it on incus; destroy removes it. In between, exec runs a command inside a machine, and snapshot / restore capture and return the whole scenario as one state.

segments as isolated links works
machines, multi-homed or detached works
declared addresses, both families works
segment MTU works
gateways, NAT, masquerade works
published: ports, as DNAT through the gateway's address works
mapping_ttl: as a conntrack timeout, read back after setting works
forwardable: false — outbound only works
policy: between segments, asymmetric works
inbound: deny as a host firewall, read back after applying works
several public networks, routed through a transit router works
diagram — the scenario drawn, from the declaration or from the hypervisor works
place: refused at raise
the lab's own certificate authority not built

What it does not do, and why that matters

place: is refused. A scenario can declare that a node host is placed on a machine; the lab names the gap and refuses rather than raising a scenario that silently lacks what it declared. Nothing can be placed because tier 0 does not exist yet.

The consequence is worth stating plainly rather than leaving to be inferred: the lab raises empty machines. It reproduces a network faithfully and puts nothing on it. It is infrastructure whose consumer has not been built, and it stays that way until tier 0 does.

The certificate story is designed and absent. 01-end-to-end-testing.md specifies the lab running its own ACME issuer on the public segment, preserving production's two-CA split. None of that is built.

What shipped differently from the design

The drawing was never designed. diagram renders a scenario as draw.io, from the declaration or from the running instance, and it exists because it was asked for during the build. It has tests and a decision record (ADR 0035, proposed) but no document in the to-be layer. It is recorded here because it runs, not because it was planned.

A router is tagged as a machine as well as a router. The design speaks of routers and machines as distinct. In the implementation a router carries user.mesh-lab.machine too, because destroy finds an instance's resources with one query and a router that carried only router= was left behind — holding its networks open, so destroy reported removing zero segments.

The router image is built once and cached. A scenario is a closed address space, so a router has no route to a package repository and cannot install nftables at raise time. The image is prepared once, with temporary connectivity. That is the only step in the whole lab that needs the workstation to be online.

The rules that turned out to be load-bearing

Public segments must use documentation ranges (RFC 5737, RFC 3849), refused by the validator before anything is raised. Research 004 found why: the mesh decides public-versus-private by matching the address, so a private range on a segment meant to be routable makes the mesh silently never form.

A scenario is a closed address space. The workstation has no route in, so two instances raised from one declaration hold the same addresses and never meet. Reachability is therefore asked from inside — exec on one machine, testing another. The workstation's opinion would be a different question with a misleadingly similar answer.

One public address is one gateway. Two gateway declarations sharing an address are one box, and their address lists union. Before this, gateways were grouped on their exact address list, and a household declaring a v6 address on one of its two segments became two router containers holding one address on one segment — which resolved to whichever answered ARP last.

What it costs

Measured on a workstation, not asserted:

one machine two machines two machines and a router
raise, to usable 12.5 s 14.6 s 32 s
snapshot 0.14 s 0.28 s —
restore, to usable again 10.5 s 11.6 s —

Machines boot concurrently, so a second machine costs seconds rather than doubling the wait. Nearly all the remaining time is boot.

These numbers depend entirely on a copy-on-write storage pool. On dir the same snapshot takes 9.9 s and a full copy of the disk, and a second did not finish in two minutes — so the lab's check refuses rather than warns. A machine without copy-on-write runs scenarios correctly and snapshots roughly 76× slower, which does not make the lab slow, it makes it unused.

How it is checked

npm run check — typecheck over source and tests, then the offline suite, then integration against a real hypervisor. Mocking the hypervisor is forbidden (ADR 0034, proposed): a test that fakes the system under integration asserts that the fake behaves as expected.

Integration tests skip with a reason on a machine that cannot raise scenarios, rather than passing green having checked nothing.

Two things the suite does not yet do, recorded because their absence is invisible:

  • It raises two of the five scenarios. Both faults found so far — two gateways holding one address, and a gateway drawn across an unrelated network — lived in scenarios nothing ever built. They were found by looking at pictures, not by running tests.
  • Nothing opens the generated draw.io file. The tests assert on the XML and check the stencil names against draw.io's own library, but no test has ever opened one.

References