Files
hq/03-DESIGN/00-as-is/11-the-lab.md
T
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00

146 lines
7.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
layer: as-is
status: implemented
code: [mesh-lab]
updated: 2026-08-25
decisions:
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
---
# The lab, as it stands
The first piece of the new shape that exists. It is the only repository outside the monorepo
so far, and unlike everything else planned it ships to nobody: it runs on a workstation,
raises virtual machines, and throws them away.
Written from the implementation. Where intent and implementation disagree, the implementation
is what is recorded here and the disagreement is stated.
## What it does
A scenario is a YAML file declaring an **underlay** — segments, the gateways between them, and
machines placed on them. `raise` materialises it on `incus`; `destroy` removes it. In between,
`exec` runs a command inside a machine, and `snapshot` / `restore` capture and return the whole
scenario as one state.
| | |
|---|---|
| segments as isolated links | works |
| machines, multi-homed or detached | works |
| declared addresses, both families | works |
| segment MTU | works |
| gateways, NAT, masquerade | works |
| `published:` ports, as DNAT through the gateway's address | works |
| `mapping_ttl:` as a conntrack timeout, read back after setting | works |
| `forwardable: false` — outbound only | works |
| `policy:` between segments, asymmetric | works |
| `inbound: deny` as a host firewall, read back after applying | works |
| several public networks, routed through a transit router | works |
| `diagram` — the scenario drawn, from the declaration or from the hypervisor | works |
| `place:` | **refused at raise** |
| the lab's own certificate authority | **not built** |
## What it does not do, and why that matters
**`place:` is refused.** A scenario can declare that a node host is placed on a machine; the
lab names the gap and refuses rather than raising a scenario that silently lacks what it
declared. Nothing can be placed because tier 0 does not exist yet.
The consequence is worth stating plainly rather than leaving to be inferred: **the lab raises
empty machines.** It reproduces a network faithfully and puts nothing on it. It is
infrastructure whose consumer has not been built, and it stays that way until tier 0 does.
**The certificate story is designed and absent.**
[`01-end-to-end-testing.md`](../01-to-be/01-end-to-end-testing.md) specifies the lab running
its own ACME issuer on the public segment, preserving production's two-CA split. None of that
is built.
## What shipped differently from the design
**The drawing was never designed.** `diagram` renders a scenario as draw.io, from the
declaration or from the running instance, and it exists because it was asked for during the
build. It has tests and a decision record ([ADR 0018](../../02-DECISIONS/0018-a-picture-is-read-from-what-runs.md),
proposed) but no document in the to-be layer. It is recorded here because it runs, not because
it was planned.
**A router is tagged as a machine as well as a router.** The design speaks of routers and
machines as distinct. In the implementation a router carries `user.mesh-lab.machine` too,
because `destroy` finds an instance's resources with one query and a router that carried only
`router=` was left behind — holding its networks open, so `destroy` reported removing zero
segments.
**The router image is built once and cached.** A scenario is a closed address space, so a
router has no route to a package repository and cannot install `nftables` at raise time. The
image is prepared once, with temporary connectivity. That is the only step in the whole lab
that needs the workstation to be online.
## The rules that turned out to be load-bearing
**Public segments must use documentation ranges** (RFC 5737, RFC 3849), refused by the
validator before anything is raised. Research 004 found why: the mesh decides
public-versus-private by matching the address, so a private range on a segment meant to be
routable makes the mesh silently never form.
**A scenario is a closed address space.** The workstation has no route in, so two instances
raised from one declaration hold the same addresses and never meet. Reachability is therefore
asked from *inside* — `exec` on one machine, testing another. The workstation's opinion would
be a different question with a misleadingly similar answer.
**One public address is one gateway.** Two gateway declarations sharing an address are one
box, and their address lists union. Before this, gateways were grouped on their exact address
list, and a household declaring a v6 address on one of its two segments became two router
containers holding one address on one segment — which resolved to whichever answered ARP last.
## What it costs
Measured on a workstation, not asserted:
| | one machine | two machines | two machines and a router |
|---|---|---|---|
| raise, to usable | 12.5 s | 14.6 s | 32 s |
| snapshot | 0.14 s | 0.28 s | — |
| restore, to usable again | 10.5 s | 11.6 s | — |
Machines boot concurrently, so a second machine costs seconds rather than doubling the wait.
Nearly all the remaining time is boot.
These numbers depend entirely on a copy-on-write storage pool. On `dir` the same snapshot takes
9.9 s and a full copy of the disk, and a second did not finish in two minutes — so the lab's
`check` **refuses** rather than warns. A machine without copy-on-write runs scenarios correctly
and snapshots roughly 76× slower, which does not make the lab slow, it makes it unused.
## How it is checked
`npm run check` — typecheck over source *and* tests, then the offline suite, then integration
against a real hypervisor. Mocking the hypervisor is forbidden
([ADR 0017](../../02-DECISIONS/0017-a-test-defends-a-decision.md), proposed): a test that fakes
the system under integration asserts that the fake behaves as expected.
Integration tests **skip with a reason** on a machine that cannot raise scenarios, rather than
passing green having checked nothing.
Two things the suite does not yet do, recorded because their absence is invisible:
- **It raises two of the five scenarios.** Both faults found so far — two gateways holding one
address, and a gateway drawn across an unrelated network — lived in scenarios nothing ever
built. They were found by looking at pictures, not by running tests.
- **Nothing opens the generated draw.io file.** The tests assert on the XML and check the
stencil names against draw.io's own library, but no test has ever opened one.
## References
- [`01-to-be/02-scenario-declaration.md`](../01-to-be/02-scenario-declaration.md) — what a
scenario declares.
- [`01-to-be/03-scenario-lifecycle.md`](../01-to-be/03-scenario-lifecycle.md) — what happens
to one.
- [`01-to-be/04-lab-installation.md`](../01-to-be/04-lab-installation.md) — what the
workstation needs.
- [Research 004](../../01-RESEARCH/004-lab-network/00-overview.md) — the topology, and the
address-range constraint the whole thing turns on.
- [Research 010](../../01-RESEARCH/010-lab-inner-loop-cost/00-overview.md) — where the measured
costs come from.