The tiers were settled and the product was named, but the repositories themselves existed only in a research sketch. That had already caused two problems. ADR 0029 makes the lab phase 0 of the migration and could not say where it lives, because no record named a repository. And the sketch contradicted an accepted record: it listed mesh-hq while ADR 0028 had decided novox/hq and explicitly rejected that name. A design resting on research is resting on something that can change without a decision. Corrected in the research too. The naming rule, which both earlier records implied and neither stated: a repository belonging to a product carries that product's prefix; a company-scoped one does not. That is why this repository is hq and the mesh's are mesh-*. Seven repositories recorded — host, substrate, control, surfaces, sdk, lab, and this one. The lab gets its own: it ships to nobody, outlives any single tier, and drives virtualisation on a workstation, which nothing else does. Inside the host it would couple development tooling to a shipped component; inside the control plane the bootstrap scenario would depend on a tier that does not exist when it is needed. Tier 4 is deliberately not decided. Whether the catalogue is one repository, one per domain or one per application stays open from ADR 0015 and is blocked on research 005 — how many repositories hold domains cannot be answered before knowing what the domains are. mesh-catalog appears in the sketch and is not decided by this record. The cost is stated rather than glossed: seven release cadences where there is one, and cross-repository changes that used to be one commit.
102 lines
5.7 KiB
Markdown
102 lines
5.7 KiB
Markdown
---
|
|
status: accepted
|
|
date: 2026-08-23
|
|
deciders: jochen
|
|
reconstructed: false
|
|
extends: 0016-a-lab-node-is-a-virtual-machine.md
|
|
---
|
|
|
|
# 29. The lab's first scenario has no pipeline, and the lab comes first
|
|
|
|
## Context
|
|
|
|
[The lab design](../03-DESIGN/01-to-be/01-end-to-end-testing.md) opens with *"what is under
|
|
test is a module; the mesh is the harness"*, and everything follows from that: a scenario has
|
|
its own forge, its own coordinator, and its own delivery cascade ending in verify. The verdict
|
|
*is* a pipeline result.
|
|
|
|
That is the right design for testing a module against the mesh that exists. It is unusable for
|
|
the thing now being built.
|
|
|
|
**The new mesh has no coordinator.** Tier 0 is a host binary and tier 1 is a pinned bundle
|
|
([research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)). A scenario that
|
|
requires a forge, a coordinator, a cascade and a meshware daemon cannot exercise them, because
|
|
all four are tier 2 and do not exist yet.
|
|
|
|
**And the sequencing was written backwards.** [Research 009](../01-RESEARCH/009-migration/00-overview.md)
|
|
placed the lab at phase B, as verification of tiers already built. But tier 0 is the component
|
|
that takes over a machine's packages, services and network — it cannot be developed against a
|
|
machine anyone needs. It needs somewhere disposable to exist **before** it is written, not
|
|
after.
|
|
|
|
## Considered options
|
|
|
|
1. **Develop tiers 0 and 1 against a real machine; add the lab afterwards.** Rejected twice
|
|
over. Developing something that reformats a machine, against a machine that is in use, is
|
|
how a machine is lost. And it would leave the bootstrap path exercised only when performed
|
|
for real — which is precisely the property that makes the current first-node script the
|
|
least-tested code in the system.
|
|
2. **Build the full lab first.** Impossible, not merely unwise: the full scenario needs a
|
|
coordinator, a forge and a delivery cascade, all of which are tier 2. It cannot precede the
|
|
tiers it is meant to test.
|
|
3. **Two scenario classes, the smaller one first, the larger a superset.** Chosen.
|
|
|
|
## Decision
|
|
|
|
The lab has **two scenario classes**, and the first has no pipeline in it at all.
|
|
|
|
| | **Bootstrap scenario** | **Full scenario** |
|
|
|---|---|---|
|
|
| Contains | one or more virtual machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules |
|
|
| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify |
|
|
| Exercises | tiers 0 and 1 | tiers 2 and above, and modules |
|
|
| Exists to | develop the mesh | test what runs on it |
|
|
|
|
The bootstrap scenario is a **strict subset** of the full one — the same virtualisation, the
|
|
same networking, the same scenario lifecycle, simply stopping before a control plane exists.
|
|
Nothing forks, which is the same rule the existing design already holds itself to.
|
|
|
|
**The lab is built first**, ahead of tier 0, and [research 009](../01-RESEARCH/009-migration/00-overview.md)
|
|
is resequenced accordingly. It is the environment everything else is developed inside.
|
|
|
|
Of the runner's two candidate jobs, this settles their order: **scenario lifecycle is needed
|
|
immediately** — something must materialise, snapshot and destroy a mesh before anything else
|
|
can be written. **Assertion execution comes later**, with the full scenario, because a
|
|
bootstrap scenario's assertions are about the state a single host reconciled and are small
|
|
enough to state directly.
|
|
|
|
## Consequences
|
|
|
|
- **The hardest path to test becomes the one exercised most.** Raising a node from nothing is
|
|
currently a script that runs when a node is created and is otherwise never touched. Under
|
|
this decision it is the inner development loop for every change to tiers 0 and 1.
|
|
- The first thing built is small: virtualisation, a network, a way to place a binary, and a way
|
|
to snapshot and reset. No forge, no coordinator, no pipeline, no modules.
|
|
- The full scenario becomes reachable by *addition* rather than by rework, because it differs
|
|
only in what is placed inside the machines.
|
|
- The lab acquires a second audience. It was designed for a module author and now also serves
|
|
whoever is building the mesh itself — which is the same "one runner, two callers" argument
|
|
the design already makes, extended one step.
|
|
- **A stale claim in the design is corrected.** It argues that scenarios are *"affordable with
|
|
system containers and would not be with virtual machines — the unit choice is what makes the
|
|
gate possible at all."* [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) superseded that:
|
|
a lab node is a virtual machine, and the scale argument for system containers was found to
|
|
have been invented rather than required. The design text did not follow the decision. It does
|
|
now.
|
|
- The lab's home is `novox/mesh-lab`, recorded in
|
|
[ADR 0030](0030-the-repository-structure.md) — written after this record, because this one
|
|
needed a repository that no decision had yet named.
|
|
- The bootstrap scenario's fidelity is its whole value, and also its risk: if it diverges from
|
|
how a real node is raised, it certifies something that does not happen. That is the same
|
|
hazard the existing design names for the full scenario, and the same answer applies —
|
|
nothing new drives it, and what runs is the real thing.
|
|
|
|
## References
|
|
|
|
- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — a lab node is a virtual machine.
|
|
- [Research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) — the tiers, and the
|
|
observation this rests on: a scenario needing only tiers 0 and 1 is one machine and a pinned
|
|
bundle, which is also exactly the bootstrap path.
|
|
- [Research 009](../01-RESEARCH/009-migration/00-overview.md) — the migration sequence this
|
|
reorders.
|