The fix is real: 76x, verified. And how the lab installs on a clean machine

Snapshot 9.9s -> 0.13s. Restore 10.4s -> 0.80s. Three snapshots sharing
1.36 GB instead of costing 4.8 GB. The projected four-machine reset cycle
falls from ~90s, unbounded at worst, to ~15s dominated by a boot that
cannot be avoided. ADR 0029's inner-loop argument holds with copy-on-write
and did not without it.

The consistency matters as much as the speed: three consecutive snapshots
took 0.13, 0.12 and 0.13 seconds, against a dir second snapshot that never
finished.

One honest counter-observation recorded: launching onto the fresh
copy-on-write pool was slower, 20.2s against 14.3s, because the image had
to be unpacked into a pool that had never seen it. Paid once per pool, and
dwarfed by what snapshotting saves, but it went the other way.

Doing the measurement produced the answer to how the lab installs on a
clean machine, because both failure modes appeared while doing it.

Installed is not available: the daemon was present with units disabled and
no group. Issue 007.

Available is not adequate, and this is worse: with the storage tooling
absent everything worked and snapshots were seventy-six times slower.
Nothing failed, nothing warned. That is a variant the mesh has not
catalogued — its usual failure is reported success and did nothing; this is
reported success and did it seventy-six times slower, which no error
surface catches because nothing is wrong.

So the lab verifies CAPABILITY, never installation, and refuses to run
degraded rather than warning — a warning about a slow inner loop is read
once and ignored forever. Prerequisites may arrive from a mesh module or
from the lab's own bootstrap, and the second path is required rather than
convenient: a lab installable only by a mesh cannot host the development
of the mesh that installs it.

The lab is the second thing installed by hand, after the node host, and for
the same reason: something has to be first, and pretending otherwise
produces a circularity papered over by a script nobody exercises.
This commit is contained in:
2026-08-24 00:14:57 +02:00
parent 98bcd5cc49
commit e88b448145
4 changed files with 173 additions and 8 deletions
@@ -37,6 +37,10 @@ Measured, and the finding is a blocker rather than a data point.
for one small virtual machine at best — and over two minutes when observed a second time. Cost
scales with the number of machines and the size of their disks, not with what changed.
**With copy-on-write it is 0.13 seconds and costs the delta.** Verified, not assumed. The
projected four-machine reset cycle falls from roughly ninety seconds to roughly fifteen, of
which almost all is a boot that cannot be avoided.
The cause is not virtual machines and not incus. It is that the host offers incus exactly one
storage driver, `dir`, which has no copy-on-write and therefore no cheap snapshot. The kernel
supports btrfs; the userspace tool that would let incus use it is simply not installed.
@@ -49,7 +53,8 @@ something the mesh already does.
| Question | Why it matters |
|---|---|
| How much does a copy-on-write pool actually improve it? Expected to be near-instant snapshots and delta-sized storage, but **expected is not measured**. | The whole inner-loop argument rests on the answer. |
| ~~How much does a copy-on-write pool actually improve it?~~ **Measured: snapshot 9.9 s → 0.13 s, restore 10.4 s → 0.80 s, three snapshots sharing 1.36 GB rather than costing 4.8 GB.** The projected four-machine cycle falls from ~90 s to ~15 s. | Answered. The inner-loop argument holds *with* copy-on-write and did not without it. |
| How does the lab install its own prerequisites on a clean machine? | The lab needs a virtualisation daemon, copy-on-write tooling and a pool before it can do anything — and it cannot depend on the mesh for them, since it is where the mesh is built. |
| Why was the second snapshot more than twelve times slower than the first? | If snapshot cost is unpredictable rather than merely high, that is worse — a loop with a variable multi-minute step is one nobody trusts. |
| Does a scenario snapshot need the machines stopped? | Stateless snapshots of a running virtual machine capture the disk but not memory. Whether a mesh restored that way is coherent is not established. |
| What is the cost at scenario scale — four machines rather than one? | Only single-machine numbers were taken. If the operation is serial, four machines is four times the wait. |
@@ -87,12 +87,55 @@ working capability — which is precisely what was just done for the virtualisat
itself, and what [`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)
is about.
## Not measured, and it matters
## The fix, measured
The copy-on-write comparison **was not run**, because running it would mean installing a package
by hand, which the mesh's rules forbid and which would have made the measurement unreproducible
anyway. Copy-on-write snapshots are expected to be near-instant with storage proportional to
what changed. That expectation is well founded and is still an expectation.
The comparison was subsequently run. One package — `btrfs-progs`, no dependencies — installed by
hand, the daemon restarted so it re-detected drivers, a copy-on-write pool created on a loop
file, and the identical image launched onto it.
Until it is measured, the correct statement is: *the current configuration is too slow, and the
likely fix is known but unverified.*
| Operation | `dir` | copy-on-write | |
|---|---|---|---|
| snapshot | 9.9 s, then **> 120 s** | **0.13 s** | ~76× faster, and *consistent* |
| snapshot again | — | 0.12 s | |
| snapshot a third time | — | 0.13 s | |
| restore call | 10.4 s | **0.80 s** | ~13× faster |
| restore, to usable | 20.1 s | **10.5 s** | the remainder is boot, which is irreducible |
| three snapshots, storage | ~4.8 GB | **1.36 GB total, shared** | cost is the delta, not the disk |
**The fix is real, and larger than expected.** Snapshot goes from ten seconds to a tenth of a
second, and — more importantly — from *wildly variable* to *flat*. Three consecutive snapshots
took 0.13, 0.12 and 0.13 seconds. On `dir` the second snapshot never finished.
Storage stops scaling with the machine and starts scaling with what changed: three snapshots of
a 1.5 GB instance occupied 1.36 GB in total, because they share.
### What it projects to
A four-machine reset-and-rerun cycle, the operation the inner loop repeats most:
| | `dir` | copy-on-write |
|---|---|---|
| snapshot the scenario | ~40 s, 6.4 GB | **~0.5 s**, delta-sized |
| restore it | ~40 s + boot | **~3 s** + boot |
| **cycle** | **~90 s, unbounded at worst** | **~15 s, dominated by boot** |
At fifteen seconds, dominated by a boot that cannot be avoided, the inner loop is viable and
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)'s argument holds.
At ninety it did not.
### One honest counter-observation
Launching onto the fresh copy-on-write pool was **slower** — 20.2 s to usable against 14.3 s —
because the image had to be unpacked into a pool that had never seen it. That cost is paid once
per pool, not per scenario, and it is dwarfed by what snapshotting saves. But it is a real
number and it went the other way.
### State this left behind
Recorded because hand-made state is exactly what the mesh's rules exist to prevent, and it must
be declared properly rather than left as an artefact of a measurement:
- `btrfs-progs` installed by hand. Its installation regenerated the boot initramfs, a side
effect worth knowing about.
- The daemon restarted once, to re-detect drivers.
- The test pool and instance were **removed**; the pool the lab actually needs does not exist.