Files
hq/02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
T
jschoubben b4904fec7e The lab comes first, and its first scenario has no pipeline
The lab was designed around a module under test, with a scenario being a
complete mesh — forge, coordinator, cascade, verify. That is unusable for
building the new mesh, because all four are tier 2 and do not exist yet.

And research 009 had the sequence backwards. It placed the lab at phase B
as verification of tiers already built, but tier 0 is the component that
takes over a machine's packages, services and network. It cannot be
developed against a machine anyone needs. The lab has to exist before the
thing it will test.

ADR 0029 splits scenarios into two classes. The bootstrap scenario is
virtual machines, the host binary and a pinned bundle, with the verdict
coming from what the host reports about the state it reconciled. The full
scenario is the designed one. The first is a strict subset of the second —
same virtualisation, same networking, same lifecycle, stopping before a
control plane exists — so the second is reached by addition rather than
rework.

The consequence worth having: raising a node from nothing stops being the
least-exercised path in the system and becomes the inner development loop.

It also settles the runner's two jobs. Scenario lifecycle is needed
immediately, because something must materialise and reset a mesh before
anything can be written against it. Assertion execution waits for the full
scenario.

Corrects a stale claim in the design while amending it: it argued
scenarios were affordable with system containers and would not be with
virtual machines. ADR 0016 superseded that reasoning and the text had not
followed.

Issue 007: the lab's first requirement is installed and unusable. The
virtualisation package is present and explicitly installed; both units are
disabled, the operator is in no group, and the client reports the server
unreachable. Not issue 001 again — that is an install failing while
reporting success. This is an install succeeding when success was not the
point. A package is files; a capability is a running service and an
identity permitted to reach it, and the module model has no vocabulary for
the second.
2026-08-23 21:57:14 +02:00

5.5 KiB

status, date, deciders, reconstructed, extends
status date deciders reconstructed extends
accepted 2026-08-23 jochen false 0016-a-lab-node-is-a-virtual-machine.md

29. The lab's first scenario has no pipeline, and the lab comes first

Context

The lab design opens with "what is under test is a module; the mesh is the harness", and everything follows from that: a scenario has its own forge, its own coordinator, and its own delivery cascade ending in verify. The verdict is a pipeline result.

That is the right design for testing a module against the mesh that exists. It is unusable for the thing now being built.

The new mesh has no coordinator. Tier 0 is a host binary and tier 1 is a pinned bundle (research 006). A scenario that requires a forge, a coordinator, a cascade and a meshware daemon cannot exercise them, because all four are tier 2 and do not exist yet.

And the sequencing was written backwards. Research 009 placed the lab at phase B, as verification of tiers already built. But tier 0 is the component that takes over a machine's packages, services and network — it cannot be developed against a machine anyone needs. It needs somewhere disposable to exist before it is written, not after.

Considered options

  1. Develop tiers 0 and 1 against a real machine; add the lab afterwards. Rejected twice over. Developing something that reformats a machine, against a machine that is in use, is how a machine is lost. And it would leave the bootstrap path exercised only when performed for real — which is precisely the property that makes the current first-node script the least-tested code in the system.
  2. Build the full lab first. Impossible, not merely unwise: the full scenario needs a coordinator, a forge and a delivery cascade, all of which are tier 2. It cannot precede the tiers it is meant to test.
  3. Two scenario classes, the smaller one first, the larger a superset. Chosen.

Decision

The lab has two scenario classes, and the first has no pipeline in it at all.

Bootstrap scenario Full scenario
Contains one or more virtual machines, the host binary, a pinned substrate bundle a complete mesh: forge, coordinator, delivery, modules
Verdict from what the host reports about the state it reconciled a pipeline result ending in verify
Exercises tiers 0 and 1 tiers 2 and above, and modules
Exists to develop the mesh test what runs on it

The bootstrap scenario is a strict subset of the full one — the same virtualisation, the same networking, the same scenario lifecycle, simply stopping before a control plane exists. Nothing forks, which is the same rule the existing design already holds itself to.

The lab is built first, ahead of tier 0, and research 009 is resequenced accordingly. It is the environment everything else is developed inside.

Of the runner's two candidate jobs, this settles their order: scenario lifecycle is needed immediately — something must materialise, snapshot and destroy a mesh before anything else can be written. Assertion execution comes later, with the full scenario, because a bootstrap scenario's assertions are about the state a single host reconciled and are small enough to state directly.

Consequences

  • The hardest path to test becomes the one exercised most. Raising a node from nothing is currently a script that runs when a node is created and is otherwise never touched. Under this decision it is the inner development loop for every change to tiers 0 and 1.
  • The first thing built is small: virtualisation, a network, a way to place a binary, and a way to snapshot and reset. No forge, no coordinator, no pipeline, no modules.
  • The full scenario becomes reachable by addition rather than by rework, because it differs only in what is placed inside the machines.
  • The lab acquires a second audience. It was designed for a module author and now also serves whoever is building the mesh itself — which is the same "one runner, two callers" argument the design already makes, extended one step.
  • A stale claim in the design is corrected. It argues that scenarios are "affordable with system containers and would not be with virtual machines — the unit choice is what makes the gate possible at all." ADR 0016 superseded that: a lab node is a virtual machine, and the scale argument for system containers was found to have been invented rather than required. The design text did not follow the decision. It does now.
  • The bootstrap scenario's fidelity is its whole value, and also its risk: if it diverges from how a real node is raised, it certifies something that does not happen. That is the same hazard the existing design names for the full scenario, and the same answer applies — nothing new drives it, and what runs is the real thing.

References

  • ADR 0016 — a lab node is a virtual machine.
  • Research 006 — the tiers, and the observation this rests on: a scenario needing only tiers 0 and 1 is one machine and a pinned bundle, which is also exactly the bootstrap path.
  • Research 009 — the migration sequence this reorders.