The lab comes first, and its first scenario has no pipeline #4

Merged
jschoubben merged 3 commits from design/lab-bootstrap-scenario into main 2026-08-23 20:24:16 +00:00
4 changed files with 209 additions and 6 deletions
Showing only changes of commit b4904fec7e - Show all commits
+6 -3
View File
@@ -30,7 +30,9 @@ Recorded because incremental is the reflex answer and it is wrong in this case.
- **Nothing external depends on it.** No users outside the operator, no service level to hold.
- **The lab exists precisely for this** ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)).
A big-bang that has been rehearsed end to end, repeatedly, on identical machines is not the
same risk as one performed for the first time on the real mesh.
same risk as one performed for the first time on the real mesh. This is also why the lab is
phase 0 rather than a verification step later: the new mesh is *developed* inside it, so by
cutover the procedure has been run hundreds of times rather than rehearsed a few.
- **Incremental would carry the rot forward.** The as-is layer documents silent failure paths,
a dead test harness and unenforced rules. A gradual migration preserves them by definition.
@@ -68,8 +70,9 @@ than a discovery.
| Phase | What | Done when |
|---|---|---|
| A | Build tier 0. The host's interface first — it carries the skeleton's biggest unproven claim. | A bare machine becomes a managed node with no mesh present. |
| B | Build tier 1 and 2. | The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts. |
| **0** | **Build the lab's bootstrap scenario** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)) — virtual machines, a network, a way to place a binary, snapshot and reset. No forge, no coordinator, no pipeline. | A machine can be raised from nothing, reset, and raised again, repeatably. |
| A | Build tier 0, **inside the lab**. The host's interface first — it carries the skeleton's biggest unproven claim. | A bare machine becomes a managed node with no mesh present. |
| B | Build tier 1 and 2. The bootstrap scenario grows into the full one by addition — the same machines, with more placed inside them. | The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts. |
| C | Enough of tier 3 to operate it. | The mesh can be driven without direct database access. |
| D | Declare the existing workloads against the new model. | The lab runs them, with copies of real data shapes. |
| E | **Rehearse the cutover in the lab** against a mesh built to resemble the real one. | Repeatable, and repeatably reversible. |
@@ -0,0 +1,98 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
extends: 0016-a-lab-node-is-a-virtual-machine.md
---
# 29. The lab's first scenario has no pipeline, and the lab comes first
## Context
[The lab design](../03-DESIGN/01-to-be/01-end-to-end-testing.md) opens with *"what is under
test is a module; the mesh is the harness"*, and everything follows from that: a scenario has
its own forge, its own coordinator, and its own delivery cascade ending in verify. The verdict
*is* a pipeline result.
That is the right design for testing a module against the mesh that exists. It is unusable for
the thing now being built.
**The new mesh has no coordinator.** Tier 0 is a host binary and tier 1 is a pinned bundle
([research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)). A scenario that
requires a forge, a coordinator, a cascade and a meshware daemon cannot exercise them, because
all four are tier 2 and do not exist yet.
**And the sequencing was written backwards.** [Research 009](../01-RESEARCH/009-migration/00-overview.md)
placed the lab at phase B, as verification of tiers already built. But tier 0 is the component
that takes over a machine's packages, services and network — it cannot be developed against a
machine anyone needs. It needs somewhere disposable to exist **before** it is written, not
after.
## Considered options
1. **Develop tiers 0 and 1 against a real machine; add the lab afterwards.** Rejected twice
over. Developing something that reformats a machine, against a machine that is in use, is
how a machine is lost. And it would leave the bootstrap path exercised only when performed
for real — which is precisely the property that makes the current first-node script the
least-tested code in the system.
2. **Build the full lab first.** Impossible, not merely unwise: the full scenario needs a
coordinator, a forge and a delivery cascade, all of which are tier 2. It cannot precede the
tiers it is meant to test.
3. **Two scenario classes, the smaller one first, the larger a superset.** Chosen.
## Decision
The lab has **two scenario classes**, and the first has no pipeline in it at all.
| | **Bootstrap scenario** | **Full scenario** |
|---|---|---|
| Contains | one or more virtual machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules |
| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify |
| Exercises | tiers 0 and 1 | tiers 2 and above, and modules |
| Exists to | develop the mesh | test what runs on it |
The bootstrap scenario is a **strict subset** of the full one — the same virtualisation, the
same networking, the same scenario lifecycle, simply stopping before a control plane exists.
Nothing forks, which is the same rule the existing design already holds itself to.
**The lab is built first**, ahead of tier 0, and [research 009](../01-RESEARCH/009-migration/00-overview.md)
is resequenced accordingly. It is the environment everything else is developed inside.
Of the runner's two candidate jobs, this settles their order: **scenario lifecycle is needed
immediately** — something must materialise, snapshot and destroy a mesh before anything else
can be written. **Assertion execution comes later**, with the full scenario, because a
bootstrap scenario's assertions are about the state a single host reconciled and are small
enough to state directly.
## Consequences
- **The hardest path to test becomes the one exercised most.** Raising a node from nothing is
currently a script that runs when a node is created and is otherwise never touched. Under
this decision it is the inner development loop for every change to tiers 0 and 1.
- The first thing built is small: virtualisation, a network, a way to place a binary, and a way
to snapshot and reset. No forge, no coordinator, no pipeline, no modules.
- The full scenario becomes reachable by *addition* rather than by rework, because it differs
only in what is placed inside the machines.
- The lab acquires a second audience. It was designed for a module author and now also serves
whoever is building the mesh itself — which is the same "one runner, two callers" argument
the design already makes, extended one step.
- **A stale claim in the design is corrected.** It argues that scenarios are *"affordable with
system containers and would not be with virtual machines — the unit choice is what makes the
gate possible at all."* [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) superseded that:
a lab node is a virtual machine, and the scale argument for system containers was found to
have been invented rather than required. The design text did not follow the decision. It does
now.
- The bootstrap scenario's fidelity is its whole value, and also its risk: if it diverges from
how a real node is raised, it certifies something that does not happen. That is the same
hazard the existing design names for the full scenario, and the same answer applies —
nothing new drives it, and what runs is the real thing.
## References
- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — a lab node is a virtual machine.
- [Research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) — the tiers, and the
observation this rests on: a scenario needing only tiers 0 and 1 is one machine and a pinned
bundle, which is also exactly the bootstrap path.
- [Research 009](../01-RESEARCH/009-migration/00-overview.md) — the migration sequence this
reorders.
+35 -3
View File
@@ -30,6 +30,36 @@ its question needs. If something can only be tested against the shape the mesh h
today, that is a gap in the vocabulary rather than a reason to privilege that shape.
## Two classes of scenario
The design below describes a scenario as a complete mesh — forge, coordinator, delivery cascade
— because what it tests is a module. **That is the larger of two classes, and not the first one
built** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)).
| | **Bootstrap scenario** | **Full scenario** |
|---|---|---|
| Contains | virtual machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules |
| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify |
| Exercises | the node host and the substrate | the control plane and everything above it |
| Exists to | **develop the mesh** | **test what runs on it** |
The bootstrap scenario is a **strict subset**: same virtualisation, same networking, same
lifecycle — it simply stops before a control plane exists. Everything from *"Where this sits in
the way work happens"* onward describes the full scenario, and applies once there is a
coordinator to describe.
**The bootstrap scenario is built first, ahead of everything it will later test.** The
component that takes over a machine's packages, services and network cannot be developed
against a machine anyone needs, and raising a node from nothing is today the least-exercised
path in the system precisely because it only ever runs for real. Making it the inner
development loop inverts that.
Of the runner's two jobs, this settles their order: **scenario lifecycle first** — something
must materialise, snapshot and destroy a mesh before anything else can be written — and
**assertion execution second**, with the full scenario, since a bootstrap scenario's assertions
concern the state one host reconciled and are small enough to state directly.
## Where this sits in the way work happens
Work reaches the mesh along one path today:
@@ -93,9 +123,11 @@ drifts.
nodes that have the module assigned. It needs to be able to run the same pipeline against
a mesh named by the request instead.
- **Scenarios must be concurrent and cheap.** Several agents working means several scenarios
at once, each needing its own network and nodes. This is affordable with system containers
and would not be with virtual machines — the unit choice is what makes the gate possible
at all.
at once, each needing its own network and nodes. A lab node is a virtual machine
([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)), and snapshots are
what make repetition cheap — restoring a scenario costs far less than building one. The
earlier argument here, that only system containers made this affordable, was superseded: the
scale it assumed was invented rather than required.
- **The gate is only as good as the verification behind it.** A module with no assertions
gets a weak gate: delivery succeeded, nothing checked. So **verification coverage becomes
the number that matters**, and it starts at approximately zero.
@@ -0,0 +1,70 @@
---
status: open
opened: 2026-08-23
located-in: [hal]
fixed-by:
amended-design:
---
# 007 — An installed package is not an available capability
## Symptom
A module declares the virtualisation package the lab needs. The package is installed —
version 7.3.0-1, recorded as explicitly installed. The client binary runs.
The capability does not exist:
| Checked | State |
|---|---|
| `incus.service` | disabled, inactive |
| `incus.socket` | disabled, inactive |
| `incus-user.socket` | disabled, inactive |
| the operator's group membership | not a member of any incus group |
| the client | reports **`Server version: unreachable`** |
Nothing failed. Nothing reported anything. The declaration was satisfied exactly as written,
and the thing it was declared for cannot be used.
## Why this is not issue 001 again
[Issue 001](../001-failed-package-install-reports-success/00-report.md) is *the install failed
and the job reported success*. This is the opposite and arguably worse: **the install
succeeded, and success was not the point.**
A package is a set of files. A capability is a running service, an enabled socket, and an
identity permitted to reach it. The module model declares the first and has no vocabulary for
the second, so the gap between them is invisible — there is no state in which the mesh believes
this node has virtualisation and is wrong, because the mesh was never asked to believe it.
The distance between the two is the same one the delivery layer already has a name for:
**transport versus effect.** A package install reports that files arrived, which is transport.
## Why it matters now
This is the first requirement of the lab
([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)), which is
phase 0 of the entire migration. The first capability the new work depends on is present,
declared, and unusable — and would have stayed unusable silently.
It also generalises. Every module that declares a package needing a unit enabled, a group
joined, a kernel module loaded, or a socket activated has this gap. Post-install work lives in
hooks, and hooks have their own recorded failure mode: thirteen were found that had never run.
## Evidence
- Package recorded as installed 2026-08-23 00:02, explicitly.
- Both units disabled and inactive; the operator in no incus group; client reports the server
unreachable.
## Open questions
- Should a module be able to declare a **capability** — a unit that must be enabled, a group
the operator must be in — rather than only the package that provides it?
- If that is what hooks are for, why is the gap invisible when a hook does not run? A hook that
never fires and a hook that fires and does nothing are indistinguishable today.
- Is this what the verify stage should be asserting? It exists, and a module's own assertions
are meant to test outcomes rather than steps — "the socket accepts a connection" is exactly
that shape.
- How many other declared packages are in this state? Nothing currently reports it, which means
the answer is unknown rather than zero.