diff --git a/01-RESEARCH/009-migration/00-overview.md b/01-RESEARCH/009-migration/00-overview.md index f72efb4..832d76a 100644 --- a/01-RESEARCH/009-migration/00-overview.md +++ b/01-RESEARCH/009-migration/00-overview.md @@ -30,7 +30,9 @@ Recorded because incremental is the reflex answer and it is wrong in this case. - **Nothing external depends on it.** No users outside the operator, no service level to hold. - **The lab exists precisely for this** ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)). A big-bang that has been rehearsed end to end, repeatedly, on identical machines is not the - same risk as one performed for the first time on the real mesh. + same risk as one performed for the first time on the real mesh. This is also why the lab is + phase 0 rather than a verification step later: the new mesh is *developed* inside it, so by + cutover the procedure has been run hundreds of times rather than rehearsed a few. - **Incremental would carry the rot forward.** The as-is layer documents silent failure paths, a dead test harness and unenforced rules. A gradual migration preserves them by definition. @@ -68,8 +70,9 @@ than a discovery. | Phase | What | Done when | |---|---|---| -| A | Build tier 0. The host's interface first — it carries the skeleton's biggest unproven claim. | A bare machine becomes a managed node with no mesh present. | -| B | Build tier 1 and 2. | The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts. | +| **0** | **Build the lab's bootstrap scenario** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)) — virtual machines, a network, a way to place a binary, snapshot and reset. No forge, no coordinator, no pipeline. | A machine can be raised from nothing, reset, and raised again, repeatably. | +| A | Build tier 0, **inside the lab**. The host's interface first — it carries the skeleton's biggest unproven claim. | A bare machine becomes a managed node with no mesh present. | +| B | Build tier 1 and 2. The bootstrap scenario grows into the full one by addition — the same machines, with more placed inside them. | The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts. | | C | Enough of tier 3 to operate it. | The mesh can be driven without direct database access. | | D | Declare the existing workloads against the new model. | The lab runs them, with copies of real data shapes. | | E | **Rehearse the cutover in the lab** against a mesh built to resemble the real one. | Repeatable, and repeatably reversible. | diff --git a/02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md b/02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md new file mode 100644 index 0000000..e1d3af6 --- /dev/null +++ b/02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md @@ -0,0 +1,98 @@ +--- +status: accepted +date: 2026-08-23 +deciders: jochen +reconstructed: false +extends: 0016-a-lab-node-is-a-virtual-machine.md +--- + +# 29. The lab's first scenario has no pipeline, and the lab comes first + +## Context + +[The lab design](../03-DESIGN/01-to-be/01-end-to-end-testing.md) opens with *"what is under +test is a module; the mesh is the harness"*, and everything follows from that: a scenario has +its own forge, its own coordinator, and its own delivery cascade ending in verify. The verdict +*is* a pipeline result. + +That is the right design for testing a module against the mesh that exists. It is unusable for +the thing now being built. + +**The new mesh has no coordinator.** Tier 0 is a host binary and tier 1 is a pinned bundle +([research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)). A scenario that +requires a forge, a coordinator, a cascade and a meshware daemon cannot exercise them, because +all four are tier 2 and do not exist yet. + +**And the sequencing was written backwards.** [Research 009](../01-RESEARCH/009-migration/00-overview.md) +placed the lab at phase B, as verification of tiers already built. But tier 0 is the component +that takes over a machine's packages, services and network — it cannot be developed against a +machine anyone needs. It needs somewhere disposable to exist **before** it is written, not +after. + +## Considered options + +1. **Develop tiers 0 and 1 against a real machine; add the lab afterwards.** Rejected twice + over. Developing something that reformats a machine, against a machine that is in use, is + how a machine is lost. And it would leave the bootstrap path exercised only when performed + for real — which is precisely the property that makes the current first-node script the + least-tested code in the system. +2. **Build the full lab first.** Impossible, not merely unwise: the full scenario needs a + coordinator, a forge and a delivery cascade, all of which are tier 2. It cannot precede the + tiers it is meant to test. +3. **Two scenario classes, the smaller one first, the larger a superset.** Chosen. + +## Decision + +The lab has **two scenario classes**, and the first has no pipeline in it at all. + +| | **Bootstrap scenario** | **Full scenario** | +|---|---|---| +| Contains | one or more virtual machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules | +| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify | +| Exercises | tiers 0 and 1 | tiers 2 and above, and modules | +| Exists to | develop the mesh | test what runs on it | + +The bootstrap scenario is a **strict subset** of the full one — the same virtualisation, the +same networking, the same scenario lifecycle, simply stopping before a control plane exists. +Nothing forks, which is the same rule the existing design already holds itself to. + +**The lab is built first**, ahead of tier 0, and [research 009](../01-RESEARCH/009-migration/00-overview.md) +is resequenced accordingly. It is the environment everything else is developed inside. + +Of the runner's two candidate jobs, this settles their order: **scenario lifecycle is needed +immediately** — something must materialise, snapshot and destroy a mesh before anything else +can be written. **Assertion execution comes later**, with the full scenario, because a +bootstrap scenario's assertions are about the state a single host reconciled and are small +enough to state directly. + +## Consequences + +- **The hardest path to test becomes the one exercised most.** Raising a node from nothing is + currently a script that runs when a node is created and is otherwise never touched. Under + this decision it is the inner development loop for every change to tiers 0 and 1. +- The first thing built is small: virtualisation, a network, a way to place a binary, and a way + to snapshot and reset. No forge, no coordinator, no pipeline, no modules. +- The full scenario becomes reachable by *addition* rather than by rework, because it differs + only in what is placed inside the machines. +- The lab acquires a second audience. It was designed for a module author and now also serves + whoever is building the mesh itself — which is the same "one runner, two callers" argument + the design already makes, extended one step. +- **A stale claim in the design is corrected.** It argues that scenarios are *"affordable with + system containers and would not be with virtual machines — the unit choice is what makes the + gate possible at all."* [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) superseded that: + a lab node is a virtual machine, and the scale argument for system containers was found to + have been invented rather than required. The design text did not follow the decision. It does + now. +- The bootstrap scenario's fidelity is its whole value, and also its risk: if it diverges from + how a real node is raised, it certifies something that does not happen. That is the same + hazard the existing design names for the full scenario, and the same answer applies — + nothing new drives it, and what runs is the real thing. + +## References + +- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — a lab node is a virtual machine. +- [Research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) — the tiers, and the + observation this rests on: a scenario needing only tiers 0 and 1 is one machine and a pinned + bundle, which is also exactly the bootstrap path. +- [Research 009](../01-RESEARCH/009-migration/00-overview.md) — the migration sequence this + reorders. diff --git a/03-DESIGN/01-to-be/01-end-to-end-testing.md b/03-DESIGN/01-to-be/01-end-to-end-testing.md index f8cdad1..b550a47 100644 --- a/03-DESIGN/01-to-be/01-end-to-end-testing.md +++ b/03-DESIGN/01-to-be/01-end-to-end-testing.md @@ -30,6 +30,36 @@ its question needs. If something can only be tested against the shape the mesh h today, that is a gap in the vocabulary rather than a reason to privilege that shape. +## Two classes of scenario + +The design below describes a scenario as a complete mesh — forge, coordinator, delivery cascade +— because what it tests is a module. **That is the larger of two classes, and not the first one +built** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)). + +| | **Bootstrap scenario** | **Full scenario** | +|---|---|---| +| Contains | virtual machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules | +| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify | +| Exercises | the node host and the substrate | the control plane and everything above it | +| Exists to | **develop the mesh** | **test what runs on it** | + +The bootstrap scenario is a **strict subset**: same virtualisation, same networking, same +lifecycle — it simply stops before a control plane exists. Everything from *"Where this sits in +the way work happens"* onward describes the full scenario, and applies once there is a +coordinator to describe. + +**The bootstrap scenario is built first, ahead of everything it will later test.** The +component that takes over a machine's packages, services and network cannot be developed +against a machine anyone needs, and raising a node from nothing is today the least-exercised +path in the system precisely because it only ever runs for real. Making it the inner +development loop inverts that. + +Of the runner's two jobs, this settles their order: **scenario lifecycle first** — something +must materialise, snapshot and destroy a mesh before anything else can be written — and +**assertion execution second**, with the full scenario, since a bootstrap scenario's assertions +concern the state one host reconciled and are small enough to state directly. + + ## Where this sits in the way work happens Work reaches the mesh along one path today: @@ -93,9 +123,11 @@ drifts. nodes that have the module assigned. It needs to be able to run the same pipeline against a mesh named by the request instead. - **Scenarios must be concurrent and cheap.** Several agents working means several scenarios - at once, each needing its own network and nodes. This is affordable with system containers - and would not be with virtual machines — the unit choice is what makes the gate possible - at all. + at once, each needing its own network and nodes. A lab node is a virtual machine + ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)), and snapshots are + what make repetition cheap — restoring a scenario costs far less than building one. The + earlier argument here, that only system containers made this affordable, was superseded: the + scale it assumed was invented rather than required. - **The gate is only as good as the verification behind it.** A module with no assertions gets a weak gate: delivery succeeded, nothing checked. So **verification coverage becomes the number that matters**, and it starts at approximately zero. diff --git a/04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md b/04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md new file mode 100644 index 0000000..ebcf3ab --- /dev/null +++ b/04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md @@ -0,0 +1,70 @@ +--- +status: open +opened: 2026-08-23 +located-in: [hal] +fixed-by: +amended-design: +--- + +# 007 — An installed package is not an available capability + +## Symptom + +A module declares the virtualisation package the lab needs. The package is installed — +version 7.3.0-1, recorded as explicitly installed. The client binary runs. + +The capability does not exist: + +| Checked | State | +|---|---| +| `incus.service` | disabled, inactive | +| `incus.socket` | disabled, inactive | +| `incus-user.socket` | disabled, inactive | +| the operator's group membership | not a member of any incus group | +| the client | reports **`Server version: unreachable`** | + +Nothing failed. Nothing reported anything. The declaration was satisfied exactly as written, +and the thing it was declared for cannot be used. + +## Why this is not issue 001 again + +[Issue 001](../001-failed-package-install-reports-success/00-report.md) is *the install failed +and the job reported success*. This is the opposite and arguably worse: **the install +succeeded, and success was not the point.** + +A package is a set of files. A capability is a running service, an enabled socket, and an +identity permitted to reach it. The module model declares the first and has no vocabulary for +the second, so the gap between them is invisible — there is no state in which the mesh believes +this node has virtualisation and is wrong, because the mesh was never asked to believe it. + +The distance between the two is the same one the delivery layer already has a name for: +**transport versus effect.** A package install reports that files arrived, which is transport. + +## Why it matters now + +This is the first requirement of the lab +([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)), which is +phase 0 of the entire migration. The first capability the new work depends on is present, +declared, and unusable — and would have stayed unusable silently. + +It also generalises. Every module that declares a package needing a unit enabled, a group +joined, a kernel module loaded, or a socket activated has this gap. Post-install work lives in +hooks, and hooks have their own recorded failure mode: thirteen were found that had never run. + +## Evidence + +- Package recorded as installed 2026-08-23 00:02, explicitly. +- Both units disabled and inactive; the operator in no incus group; client reports the server + unreachable. + +## Open questions + +- Should a module be able to declare a **capability** — a unit that must be enabled, a group + the operator must be in — rather than only the package that provides it? +- If that is what hooks are for, why is the gap invisible when a hook does not run? A hook that + never fires and a hook that fires and does nothing are indistinguishable today. +- Is this what the verify stage should be asserting? It exists, and a module's own assertions + are meant to test outcomes rather than steps — "the socket accepts a connection" is exactly + that shape. +- How many other declared packages are in this state? Nothing currently reports it, which means + the answer is unknown rather than zero.