The lab comes first, and its first scenario has no pipeline

The lab was designed around a module under test, with a scenario being a
complete mesh — forge, coordinator, cascade, verify. That is unusable for
building the new mesh, because all four are tier 2 and do not exist yet.

And research 009 had the sequence backwards. It placed the lab at phase B
as verification of tiers already built, but tier 0 is the component that
takes over a machine's packages, services and network. It cannot be
developed against a machine anyone needs. The lab has to exist before the
thing it will test.

ADR 0029 splits scenarios into two classes. The bootstrap scenario is
virtual machines, the host binary and a pinned bundle, with the verdict
coming from what the host reports about the state it reconciled. The full
scenario is the designed one. The first is a strict subset of the second —
same virtualisation, same networking, same lifecycle, stopping before a
control plane exists — so the second is reached by addition rather than
rework.

The consequence worth having: raising a node from nothing stops being the
least-exercised path in the system and becomes the inner development loop.

It also settles the runner's two jobs. Scenario lifecycle is needed
immediately, because something must materialise and reset a mesh before
anything can be written against it. Assertion execution waits for the full
scenario.

Corrects a stale claim in the design while amending it: it argued
scenarios were affordable with system containers and would not be with
virtual machines. ADR 0016 superseded that reasoning and the text had not
followed.

Issue 007: the lab's first requirement is installed and unusable. The
virtualisation package is present and explicitly installed; both units are
disabled, the operator is in no group, and the client reports the server
unreachable. Not issue 001 again — that is an install failing while
reporting success. This is an install succeeding when success was not the
point. A package is files; a capability is a running service and an
identity permitted to reach it, and the module model has no vocabulary for
the second.
This commit is contained in:
2026-08-23 21:57:14 +02:00
parent 1570234ac0
commit b4904fec7e
4 changed files with 209 additions and 6 deletions
+35 -3
View File
@@ -30,6 +30,36 @@ its question needs. If something can only be tested against the shape the mesh h
today, that is a gap in the vocabulary rather than a reason to privilege that shape.
## Two classes of scenario
The design below describes a scenario as a complete mesh — forge, coordinator, delivery cascade
— because what it tests is a module. **That is the larger of two classes, and not the first one
built** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)).
| | **Bootstrap scenario** | **Full scenario** |
|---|---|---|
| Contains | virtual machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules |
| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify |
| Exercises | the node host and the substrate | the control plane and everything above it |
| Exists to | **develop the mesh** | **test what runs on it** |
The bootstrap scenario is a **strict subset**: same virtualisation, same networking, same
lifecycle — it simply stops before a control plane exists. Everything from *"Where this sits in
the way work happens"* onward describes the full scenario, and applies once there is a
coordinator to describe.
**The bootstrap scenario is built first, ahead of everything it will later test.** The
component that takes over a machine's packages, services and network cannot be developed
against a machine anyone needs, and raising a node from nothing is today the least-exercised
path in the system precisely because it only ever runs for real. Making it the inner
development loop inverts that.
Of the runner's two jobs, this settles their order: **scenario lifecycle first** — something
must materialise, snapshot and destroy a mesh before anything else can be written — and
**assertion execution second**, with the full scenario, since a bootstrap scenario's assertions
concern the state one host reconciled and are small enough to state directly.
## Where this sits in the way work happens
Work reaches the mesh along one path today:
@@ -93,9 +123,11 @@ drifts.
nodes that have the module assigned. It needs to be able to run the same pipeline against
a mesh named by the request instead.
- **Scenarios must be concurrent and cheap.** Several agents working means several scenarios
at once, each needing its own network and nodes. This is affordable with system containers
and would not be with virtual machines — the unit choice is what makes the gate possible
at all.
at once, each needing its own network and nodes. A lab node is a virtual machine
([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)), and snapshots are
what make repetition cheap — restoring a scenario costs far less than building one. The
earlier argument here, that only system containers made this affordable, was superseded: the
scale it assumed was invented rather than required.
- **The gate is only as good as the verification behind it.** A module with no assertions
gets a weak gate: delivery succeeded, nothing checked. So **verification coverage becomes
the number that matters**, and it starts at approximately zero.