diff --git a/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md b/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md index bd52bde..db40f90 100644 --- a/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md +++ b/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md @@ -37,6 +37,10 @@ Measured, and the finding is a blocker rather than a data point. for one small virtual machine at best — and over two minutes when observed a second time. Cost scales with the number of machines and the size of their disks, not with what changed. +**With copy-on-write it is 0.13 seconds and costs the delta.** Verified, not assumed. The +projected four-machine reset cycle falls from roughly ninety seconds to roughly fifteen, of +which almost all is a boot that cannot be avoided. + The cause is not virtual machines and not incus. It is that the host offers incus exactly one storage driver, `dir`, which has no copy-on-write and therefore no cheap snapshot. The kernel supports btrfs; the userspace tool that would let incus use it is simply not installed. @@ -49,7 +53,8 @@ something the mesh already does. | Question | Why it matters | |---|---| -| How much does a copy-on-write pool actually improve it? Expected to be near-instant snapshots and delta-sized storage, but **expected is not measured**. | The whole inner-loop argument rests on the answer. | +| ~~How much does a copy-on-write pool actually improve it?~~ **Measured: snapshot 9.9 s → 0.13 s, restore 10.4 s → 0.80 s, three snapshots sharing 1.36 GB rather than costing 4.8 GB.** The projected four-machine cycle falls from ~90 s to ~15 s. | Answered. The inner-loop argument holds *with* copy-on-write and did not without it. | +| How does the lab install its own prerequisites on a clean machine? | The lab needs a virtualisation daemon, copy-on-write tooling and a pool before it can do anything — and it cannot depend on the mesh for them, since it is where the mesh is built. | | Why was the second snapshot more than twelve times slower than the first? | If snapshot cost is unpredictable rather than merely high, that is worse — a loop with a variable multi-minute step is one nobody trusts. | | Does a scenario snapshot need the machines stopped? | Stateless snapshots of a running virtual machine capture the disk but not memory. Whether a mesh restored that way is coherent is not established. | | What is the cost at scenario scale — four machines rather than one? | Only single-machine numbers were taken. If the operation is serial, four machines is four times the wait. | diff --git a/01-RESEARCH/010-lab-inner-loop-cost/measurements.md b/01-RESEARCH/010-lab-inner-loop-cost/measurements.md index 9189c07..a690054 100644 --- a/01-RESEARCH/010-lab-inner-loop-cost/measurements.md +++ b/01-RESEARCH/010-lab-inner-loop-cost/measurements.md @@ -87,12 +87,55 @@ working capability — which is precisely what was just done for the virtualisat itself, and what [`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md) is about. -## Not measured, and it matters +## The fix, measured -The copy-on-write comparison **was not run**, because running it would mean installing a package -by hand, which the mesh's rules forbid and which would have made the measurement unreproducible -anyway. Copy-on-write snapshots are expected to be near-instant with storage proportional to -what changed. That expectation is well founded and is still an expectation. +The comparison was subsequently run. One package — `btrfs-progs`, no dependencies — installed by +hand, the daemon restarted so it re-detected drivers, a copy-on-write pool created on a loop +file, and the identical image launched onto it. -Until it is measured, the correct statement is: *the current configuration is too slow, and the -likely fix is known but unverified.* +| Operation | `dir` | copy-on-write | | +|---|---|---|---| +| snapshot | 9.9 s, then **> 120 s** | **0.13 s** | ~76× faster, and *consistent* | +| snapshot again | — | 0.12 s | | +| snapshot a third time | — | 0.13 s | | +| restore call | 10.4 s | **0.80 s** | ~13× faster | +| restore, to usable | 20.1 s | **10.5 s** | the remainder is boot, which is irreducible | +| three snapshots, storage | ~4.8 GB | **1.36 GB total, shared** | cost is the delta, not the disk | + +**The fix is real, and larger than expected.** Snapshot goes from ten seconds to a tenth of a +second, and — more importantly — from *wildly variable* to *flat*. Three consecutive snapshots +took 0.13, 0.12 and 0.13 seconds. On `dir` the second snapshot never finished. + +Storage stops scaling with the machine and starts scaling with what changed: three snapshots of +a 1.5 GB instance occupied 1.36 GB in total, because they share. + +### What it projects to + +A four-machine reset-and-rerun cycle, the operation the inner loop repeats most: + +| | `dir` | copy-on-write | +|---|---|---| +| snapshot the scenario | ~40 s, 6.4 GB | **~0.5 s**, delta-sized | +| restore it | ~40 s + boot | **~3 s** + boot | +| **cycle** | **~90 s, unbounded at worst** | **~15 s, dominated by boot** | + +At fifteen seconds, dominated by a boot that cannot be avoided, the inner loop is viable and +[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)'s argument holds. +At ninety it did not. + +### One honest counter-observation + +Launching onto the fresh copy-on-write pool was **slower** — 20.2 s to usable against 14.3 s — +because the image had to be unpacked into a pool that had never seen it. That cost is paid once +per pool, not per scenario, and it is dwarfed by what snapshotting saves. But it is a real +number and it went the other way. + +### State this left behind + +Recorded because hand-made state is exactly what the mesh's rules exist to prevent, and it must +be declared properly rather than left as an artefact of a measurement: + +- `btrfs-progs` installed by hand. Its installation regenerated the boot initramfs, a side + effect worth knowing about. +- The daemon restarted once, to re-detect drivers. +- The test pool and instance were **removed**; the pool the lab actually needs does not exist. diff --git a/03-DESIGN/01-to-be/04-lab-installation.md b/03-DESIGN/01-to-be/04-lab-installation.md new file mode 100644 index 0000000..c1bd0cd --- /dev/null +++ b/03-DESIGN/01-to-be/04-lab-installation.md @@ -0,0 +1,116 @@ +--- +layer: to-be +status: designed +code: [mesh-lab] +updated: 2026-08-24 +decisions: + - 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md + - 02-DECISIONS/0008-a-failed-step-fails-the-job.md +--- + +# Installing the lab on a clean machine + +The lab has prerequisites — a virtualisation daemon, copy-on-write storage, a pool, an identity +permitted to talk to it — and it cannot get them from the mesh, because it is where the mesh is +built ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)). + +So the lab needs an install path of its own. This describes it, and the shape it has to take is +determined by two failures observed while measuring +([research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md)). + +## The two failures that shape this + +**One: installed is not available.** The virtualisation package was present and explicitly +installed. Both its units were disabled, the operator was in no group, and the client reported +the server unreachable. Nothing had failed — the declaration was satisfied exactly as written +([`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)). + +**Two, and worse: available is not adequate.** With the storage tooling absent, the daemon +offered one driver, everything worked, and snapshots took **seventy-six times longer** than they +needed to. Nothing failed. Nothing warned. A lab in that state runs correctly and is simply too +slow to use — and the inner loop it exists to provide quietly does not exist. + +The second is the more dangerous shape and it is a variant the mesh has not catalogued before. +Its usual failure is *reported success and did nothing*. This is **reported success and did it +seventy-six times slower**, which no error surface catches because nothing is wrong. + +## What follows: the lab verifies capability, never installation + +The install path may differ. **The verification does not.** + +Before the lab will raise anything, it asserts the outcomes it needs — not that packages are +present, but that the machine can actually do the work: + +| Assertion | Failing means | +|---|---| +| the daemon answers **as the invoking user**, not as root | a group membership that was granted but never took effect | +| a copy-on-write storage driver is offered | the userspace tooling is missing; snapshots will be full copies | +| **the pool the lab will use is on that driver** | a pool exists, and is the slow kind — the failure that has no symptom | +| hardware virtualisation is present | machines will be emulated and unusably slow | +| an image can be fetched or is cached | the first raise will fail late instead of early | + +Each check states **why it matters**, in the terms of what it costs. *"The pool uses the `dir` +driver"* means nothing to someone who does not already know it means seventy-six times slower +and unbounded at worst. + +**The lab refuses to run degraded.** It does not warn and continue: a warning about a slow inner +loop is read once and ignored forever, and the loop stays slow. This is +[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) applied where the failure is +performance rather than an error. + +## Two ways the prerequisites arrive + +**On a machine the mesh manages** — a module declares them, and a hook turns them into +capabilities: the units enabled, the group granted, the pool created on the right driver. This +already works; it is what was done for the virtualisation daemon itself. + +**On a machine the mesh does not manage** — the lab's own bootstrap does it. One command, on a +clean machine, that installs what is missing and configures it. + +The second path is not a convenience. It is **required**, because the lab must work before the +mesh does, and a lab that could only be installed by a mesh would be unable to host the +development of the mesh that installs it. + +Both paths end at the same verification. Whoever satisfied the prerequisites, the lab checks +them itself — because the lesson of the first failure is precisely that *something else said it +was done* is not evidence. + +## The lab is the second thing installed by hand + +Worth stating, because it looks like an exception and is not. + +The node host is the one thing installed by hand on a machine +([research 006](../../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)): everything else +arrives through it. The lab is the same shape on a workstation — installed once, by hand, and +then everything about the mesh is developed inside it. + +Two bootstraps, at two levels, for the same reason: **something has to be first, and pretending +otherwise produces a circularity that gets papered over with a script nobody exercises.** + +## What a clean install actually needs + +In order, on a machine with nothing: + +1. **A virtualisation daemon**, running, with its socket enabled. +2. **Copy-on-write storage tooling** — the kernel side is usually already present; it is the + userspace half that is missing and that decides whether the driver is offered at all. +3. **A pool on that driver.** A loop-backed file is sufficient and needs no partitioning, which + matters: requiring a dedicated filesystem would make the lab uninstallable on a machine + already in use. +4. **Group membership** for the operator — which does **not** apply to sessions that already + existed. Observed directly: a shell whose process tree predated the grant could not reach + the daemon while a fresh lookup showed the membership present. The bootstrap has to say so, + or the first thing a person meets is a permission error that looks like a broken install. +5. **Verification**, as above, before anything is raised. + +## Open + +- **Whether the lab's bootstrap may install packages at all**, given that the mesh's rules + forbid installing by hand. The resolution is probably that the lab's bootstrap *is* the + sanctioned mechanism on an unmanaged machine, in the way the mesh's own first-node script is — + but that is an argument to record, not to assume. +- **What "adequate" means numerically.** The checks above are qualitative. A snapshot-time + threshold would catch a copy-on-write pool that is slow for some other reason, and would be a + real assertion rather than a proxy. +- **Whether the lab should own its pool** rather than using an existing one. Owning it makes the + driver guaranteed; sharing it avoids duplicating storage on a machine that already has a pool. diff --git a/03-DESIGN/01-to-be/README.md b/03-DESIGN/01-to-be/README.md index 70e82bf..7f2fcc7 100644 --- a/03-DESIGN/01-to-be/README.md +++ b/03-DESIGN/01-to-be/README.md @@ -13,6 +13,7 @@ document is written and this one's status becomes `implemented`. | [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md), [0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) | | [`02-scenario-declaration.md`](02-scenario-declaration.md) | What a scenario declares — the underlay, and what to place on it | [ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md) | | [`03-scenario-lifecycle.md`](03-scenario-lifecycle.md) | What happens to a scenario — raise, snapshot, restore, move, destroy | [ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md) | +| [`04-lab-installation.md`](04-lab-installation.md) | Getting the lab onto a clean machine, and why it verifies capability rather than installation | [ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) | ## Not yet written