The lab comes first, and its first scenario has no pipeline #4
@@ -30,7 +30,9 @@ Recorded because incremental is the reflex answer and it is wrong in this case.
|
||||
- **Nothing external depends on it.** No users outside the operator, no service level to hold.
|
||||
- **The lab exists precisely for this** ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)).
|
||||
A big-bang that has been rehearsed end to end, repeatedly, on identical machines is not the
|
||||
same risk as one performed for the first time on the real mesh.
|
||||
same risk as one performed for the first time on the real mesh. This is also why the lab is
|
||||
phase 0 rather than a verification step later: the new mesh is *developed* inside it, so by
|
||||
cutover the procedure has been run hundreds of times rather than rehearsed a few.
|
||||
- **Incremental would carry the rot forward.** The as-is layer documents silent failure paths,
|
||||
a dead test harness and unenforced rules. A gradual migration preserves them by definition.
|
||||
|
||||
@@ -68,8 +70,9 @@ than a discovery.
|
||||
|
||||
| Phase | What | Done when |
|
||||
|---|---|---|
|
||||
| A | Build tier 0. The host's interface first — it carries the skeleton's biggest unproven claim. | A bare machine becomes a managed node with no mesh present. |
|
||||
| B | Build tier 1 and 2. | The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts. |
|
||||
| **0** | **Build the lab's bootstrap scenario** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)) — virtual machines, a network, a way to place a binary, snapshot and reset. No forge, no coordinator, no pipeline. | A machine can be raised from nothing, reset, and raised again, repeatably. |
|
||||
| A | Build tier 0, **inside the lab**. The host's interface first — it carries the skeleton's biggest unproven claim. | A bare machine becomes a managed node with no mesh present. |
|
||||
| B | Build tier 1 and 2. The bootstrap scenario grows into the full one by addition — the same machines, with more placed inside them. | The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts. |
|
||||
| C | Enough of tier 3 to operate it. | The mesh can be driven without direct database access. |
|
||||
| D | Declare the existing workloads against the new model. | The lab runs them, with copies of real data shapes. |
|
||||
| E | **Rehearse the cutover in the lab** against a mesh built to resemble the real one. | Repeatable, and repeatably reversible. |
|
||||
|
||||
@@ -0,0 +1,98 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-08-23
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 0016-a-lab-node-is-a-virtual-machine.md
|
||||
---
|
||||
|
||||
# 29. The lab's first scenario has no pipeline, and the lab comes first
|
||||
|
||||
## Context
|
||||
|
||||
[The lab design](../03-DESIGN/01-to-be/01-end-to-end-testing.md) opens with *"what is under
|
||||
test is a module; the mesh is the harness"*, and everything follows from that: a scenario has
|
||||
its own forge, its own coordinator, and its own delivery cascade ending in verify. The verdict
|
||||
*is* a pipeline result.
|
||||
|
||||
That is the right design for testing a module against the mesh that exists. It is unusable for
|
||||
the thing now being built.
|
||||
|
||||
**The new mesh has no coordinator.** Tier 0 is a host binary and tier 1 is a pinned bundle
|
||||
([research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)). A scenario that
|
||||
requires a forge, a coordinator, a cascade and a meshware daemon cannot exercise them, because
|
||||
all four are tier 2 and do not exist yet.
|
||||
|
||||
**And the sequencing was written backwards.** [Research 009](../01-RESEARCH/009-migration/00-overview.md)
|
||||
placed the lab at phase B, as verification of tiers already built. But tier 0 is the component
|
||||
that takes over a machine's packages, services and network — it cannot be developed against a
|
||||
machine anyone needs. It needs somewhere disposable to exist **before** it is written, not
|
||||
after.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Develop tiers 0 and 1 against a real machine; add the lab afterwards.** Rejected twice
|
||||
over. Developing something that reformats a machine, against a machine that is in use, is
|
||||
how a machine is lost. And it would leave the bootstrap path exercised only when performed
|
||||
for real — which is precisely the property that makes the current first-node script the
|
||||
least-tested code in the system.
|
||||
2. **Build the full lab first.** Impossible, not merely unwise: the full scenario needs a
|
||||
coordinator, a forge and a delivery cascade, all of which are tier 2. It cannot precede the
|
||||
tiers it is meant to test.
|
||||
3. **Two scenario classes, the smaller one first, the larger a superset.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
The lab has **two scenario classes**, and the first has no pipeline in it at all.
|
||||
|
||||
| | **Bootstrap scenario** | **Full scenario** |
|
||||
|---|---|---|
|
||||
| Contains | one or more virtual machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules |
|
||||
| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify |
|
||||
| Exercises | tiers 0 and 1 | tiers 2 and above, and modules |
|
||||
| Exists to | develop the mesh | test what runs on it |
|
||||
|
||||
The bootstrap scenario is a **strict subset** of the full one — the same virtualisation, the
|
||||
same networking, the same scenario lifecycle, simply stopping before a control plane exists.
|
||||
Nothing forks, which is the same rule the existing design already holds itself to.
|
||||
|
||||
**The lab is built first**, ahead of tier 0, and [research 009](../01-RESEARCH/009-migration/00-overview.md)
|
||||
is resequenced accordingly. It is the environment everything else is developed inside.
|
||||
|
||||
Of the runner's two candidate jobs, this settles their order: **scenario lifecycle is needed
|
||||
immediately** — something must materialise, snapshot and destroy a mesh before anything else
|
||||
can be written. **Assertion execution comes later**, with the full scenario, because a
|
||||
bootstrap scenario's assertions are about the state a single host reconciled and are small
|
||||
enough to state directly.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **The hardest path to test becomes the one exercised most.** Raising a node from nothing is
|
||||
currently a script that runs when a node is created and is otherwise never touched. Under
|
||||
this decision it is the inner development loop for every change to tiers 0 and 1.
|
||||
- The first thing built is small: virtualisation, a network, a way to place a binary, and a way
|
||||
to snapshot and reset. No forge, no coordinator, no pipeline, no modules.
|
||||
- The full scenario becomes reachable by *addition* rather than by rework, because it differs
|
||||
only in what is placed inside the machines.
|
||||
- The lab acquires a second audience. It was designed for a module author and now also serves
|
||||
whoever is building the mesh itself — which is the same "one runner, two callers" argument
|
||||
the design already makes, extended one step.
|
||||
- **A stale claim in the design is corrected.** It argues that scenarios are *"affordable with
|
||||
system containers and would not be with virtual machines — the unit choice is what makes the
|
||||
gate possible at all."* [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) superseded that:
|
||||
a lab node is a virtual machine, and the scale argument for system containers was found to
|
||||
have been invented rather than required. The design text did not follow the decision. It does
|
||||
now.
|
||||
- The bootstrap scenario's fidelity is its whole value, and also its risk: if it diverges from
|
||||
how a real node is raised, it certifies something that does not happen. That is the same
|
||||
hazard the existing design names for the full scenario, and the same answer applies —
|
||||
nothing new drives it, and what runs is the real thing.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — a lab node is a virtual machine.
|
||||
- [Research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) — the tiers, and the
|
||||
observation this rests on: a scenario needing only tiers 0 and 1 is one machine and a pinned
|
||||
bundle, which is also exactly the bootstrap path.
|
||||
- [Research 009](../01-RESEARCH/009-migration/00-overview.md) — the migration sequence this
|
||||
reorders.
|
||||
@@ -30,6 +30,36 @@ its question needs. If something can only be tested against the shape the mesh h
|
||||
today, that is a gap in the vocabulary rather than a reason to privilege that shape.
|
||||
|
||||
|
||||
## Two classes of scenario
|
||||
|
||||
The design below describes a scenario as a complete mesh — forge, coordinator, delivery cascade
|
||||
— because what it tests is a module. **That is the larger of two classes, and not the first one
|
||||
built** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)).
|
||||
|
||||
| | **Bootstrap scenario** | **Full scenario** |
|
||||
|---|---|---|
|
||||
| Contains | virtual machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules |
|
||||
| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify |
|
||||
| Exercises | the node host and the substrate | the control plane and everything above it |
|
||||
| Exists to | **develop the mesh** | **test what runs on it** |
|
||||
|
||||
The bootstrap scenario is a **strict subset**: same virtualisation, same networking, same
|
||||
lifecycle — it simply stops before a control plane exists. Everything from *"Where this sits in
|
||||
the way work happens"* onward describes the full scenario, and applies once there is a
|
||||
coordinator to describe.
|
||||
|
||||
**The bootstrap scenario is built first, ahead of everything it will later test.** The
|
||||
component that takes over a machine's packages, services and network cannot be developed
|
||||
against a machine anyone needs, and raising a node from nothing is today the least-exercised
|
||||
path in the system precisely because it only ever runs for real. Making it the inner
|
||||
development loop inverts that.
|
||||
|
||||
Of the runner's two jobs, this settles their order: **scenario lifecycle first** — something
|
||||
must materialise, snapshot and destroy a mesh before anything else can be written — and
|
||||
**assertion execution second**, with the full scenario, since a bootstrap scenario's assertions
|
||||
concern the state one host reconciled and are small enough to state directly.
|
||||
|
||||
|
||||
## Where this sits in the way work happens
|
||||
|
||||
Work reaches the mesh along one path today:
|
||||
@@ -93,9 +123,11 @@ drifts.
|
||||
nodes that have the module assigned. It needs to be able to run the same pipeline against
|
||||
a mesh named by the request instead.
|
||||
- **Scenarios must be concurrent and cheap.** Several agents working means several scenarios
|
||||
at once, each needing its own network and nodes. This is affordable with system containers
|
||||
and would not be with virtual machines — the unit choice is what makes the gate possible
|
||||
at all.
|
||||
at once, each needing its own network and nodes. A lab node is a virtual machine
|
||||
([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)), and snapshots are
|
||||
what make repetition cheap — restoring a scenario costs far less than building one. The
|
||||
earlier argument here, that only system containers made this affordable, was superseded: the
|
||||
scale it assumed was invented rather than required.
|
||||
- **The gate is only as good as the verification behind it.** A module with no assertions
|
||||
gets a weak gate: delivery succeeded, nothing checked. So **verification coverage becomes
|
||||
the number that matters**, and it starts at approximately zero.
|
||||
|
||||
@@ -0,0 +1,70 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-08-23
|
||||
located-in: [hal]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 007 — An installed package is not an available capability
|
||||
|
||||
## Symptom
|
||||
|
||||
A module declares the virtualisation package the lab needs. The package is installed —
|
||||
version 7.3.0-1, recorded as explicitly installed. The client binary runs.
|
||||
|
||||
The capability does not exist:
|
||||
|
||||
| Checked | State |
|
||||
|---|---|
|
||||
| `incus.service` | disabled, inactive |
|
||||
| `incus.socket` | disabled, inactive |
|
||||
| `incus-user.socket` | disabled, inactive |
|
||||
| the operator's group membership | not a member of any incus group |
|
||||
| the client | reports **`Server version: unreachable`** |
|
||||
|
||||
Nothing failed. Nothing reported anything. The declaration was satisfied exactly as written,
|
||||
and the thing it was declared for cannot be used.
|
||||
|
||||
## Why this is not issue 001 again
|
||||
|
||||
[Issue 001](../001-failed-package-install-reports-success/00-report.md) is *the install failed
|
||||
and the job reported success*. This is the opposite and arguably worse: **the install
|
||||
succeeded, and success was not the point.**
|
||||
|
||||
A package is a set of files. A capability is a running service, an enabled socket, and an
|
||||
identity permitted to reach it. The module model declares the first and has no vocabulary for
|
||||
the second, so the gap between them is invisible — there is no state in which the mesh believes
|
||||
this node has virtualisation and is wrong, because the mesh was never asked to believe it.
|
||||
|
||||
The distance between the two is the same one the delivery layer already has a name for:
|
||||
**transport versus effect.** A package install reports that files arrived, which is transport.
|
||||
|
||||
## Why it matters now
|
||||
|
||||
This is the first requirement of the lab
|
||||
([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)), which is
|
||||
phase 0 of the entire migration. The first capability the new work depends on is present,
|
||||
declared, and unusable — and would have stayed unusable silently.
|
||||
|
||||
It also generalises. Every module that declares a package needing a unit enabled, a group
|
||||
joined, a kernel module loaded, or a socket activated has this gap. Post-install work lives in
|
||||
hooks, and hooks have their own recorded failure mode: thirteen were found that had never run.
|
||||
|
||||
## Evidence
|
||||
|
||||
- Package recorded as installed 2026-08-23 00:02, explicitly.
|
||||
- Both units disabled and inactive; the operator in no incus group; client reports the server
|
||||
unreachable.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should a module be able to declare a **capability** — a unit that must be enabled, a group
|
||||
the operator must be in — rather than only the package that provides it?
|
||||
- If that is what hooks are for, why is the gap invisible when a hook does not run? A hook that
|
||||
never fires and a hook that fires and does nothing are indistinguishable today.
|
||||
- Is this what the verify stage should be asserting? It exists, and a module's own assertions
|
||||
are meant to test outcomes rather than steps — "the socket accepts a connection" is exactly
|
||||
that shape.
|
||||
- How many other declared packages are in this state? Nothing currently reports it, which means
|
||||
the answer is unknown rather than zero.
|
||||
Reference in New Issue
Block a user