Files
hq/01-RESEARCH/010-lab-inner-loop-cost/measurements.md
T
jschoubben e88b448145 The fix is real: 76x, verified. And how the lab installs on a clean machine
Snapshot 9.9s -> 0.13s. Restore 10.4s -> 0.80s. Three snapshots sharing
1.36 GB instead of costing 4.8 GB. The projected four-machine reset cycle
falls from ~90s, unbounded at worst, to ~15s dominated by a boot that
cannot be avoided. ADR 0029's inner-loop argument holds with copy-on-write
and did not without it.

The consistency matters as much as the speed: three consecutive snapshots
took 0.13, 0.12 and 0.13 seconds, against a dir second snapshot that never
finished.

One honest counter-observation recorded: launching onto the fresh
copy-on-write pool was slower, 20.2s against 14.3s, because the image had
to be unpacked into a pool that had never seen it. Paid once per pool, and
dwarfed by what snapshotting saves, but it went the other way.

Doing the measurement produced the answer to how the lab installs on a
clean machine, because both failure modes appeared while doing it.

Installed is not available: the daemon was present with units disabled and
no group. Issue 007.

Available is not adequate, and this is worse: with the storage tooling
absent everything worked and snapshots were seventy-six times slower.
Nothing failed, nothing warned. That is a variant the mesh has not
catalogued — its usual failure is reported success and did nothing; this is
reported success and did it seventy-six times slower, which no error
surface catches because nothing is wrong.

So the lab verifies CAPABILITY, never installation, and refuses to run
degraded rather than warning — a warning about a slow inner loop is read
once and ignored forever. Prerequisites may arrive from a mesh module or
from the lab's own bootstrap, and the second path is required rather than
convenient: a lab installable only by a mesh cannot host the development
of the mesh that installs it.

The lab is the second thing installed by hand, after the node host, and for
the same reason: something has to be first, and pretending otherwise
produces a circularity papered over by a script nobody exercises.
2026-08-24 00:14:57 +02:00

142 lines
6.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
effort: 010-lab-inner-loop-cost
updated: 2026-08-24
---
# Measurements
Taken 2026-08-24 on a workstation with hardware virtualisation available, an NVMe-backed ext4
root, and 300 GB free. One virtual machine, 1 GiB memory, 2 CPUs, from a cached distribution
image.
## The environment, before anything ran
| Fact | Value | Consequence |
|---|---|---|
| Hardware virtualisation | present | virtual machines run at native speed; the choice in [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md) is not paying an emulation penalty |
| Storage drivers the daemon offers | **`dir` only** | no copy-on-write, therefore no cheap snapshot |
| Host filesystems | ext4 throughout | nothing copy-on-write to put a pool on |
| btrfs kernel module | **available** | the kernel can do it |
| `btrfs-progs` | **not installed** | which is the entire reason the driver is absent |
The last two rows are the finding. The daemon advertises only `dir` because the userspace tool
for anything better is missing — not because the host cannot do better.
## Raising a machine
| Step | Time |
|---|---|
| launch call returns | **3.4 s** |
| machine actually usable — a command executes on it | **14.3 s** |
The gap matters for the lifecycle design: `raise` returning is not the same as the scenario
being ready, so the verb has to wait for the second number, not report the first. Reporting
the first would be the mesh's own recurring failure — transport reported as effect.
## Snapshot and restore
| Operation | Time | Disk |
|---|---|---|
| snapshot, first | **9.9 s** | **+1.6 GB** |
| snapshot, second | **> 120 s — did not complete** | — |
| restore call returns | **10.4 s** | — |
| machine usable again | **20.1 s** total | — |
Instance on disk before snapshotting: 1.5 GB. Snapshot directory afterwards: 1.6 GB. **A `dir`
snapshot is a full copy** — the storage cost equals the instance, and nothing is shared.
Implied copy throughput on the first snapshot is roughly 160 MB/s, which is far below what the
underlying NVMe can do and is consistent with a real, durable copy rather than a metadata
operation.
**The second snapshot is the more troubling number.** It exceeded two minutes and was still
running when the observation was cut off; only the first snapshot exists. Whatever the cause —
page cache exhausted by the preceding restore, writeback contention — the practical
consequence is that snapshot cost here is **not merely high, it is unpredictable**.
## What this projects to
A four-machine scenario, taking the optimistic single-machine numbers and assuming the
operations are serial:
| | one machine | four machines |
|---|---|---|
| raise, to usable | 14 s | ~57 s |
| snapshot | 10 s, 1.6 GB | ~40 s, 6.4 GB |
| restore, to usable | 20 s | ~80 s |
A reset-and-rerun cycle is therefore **around a minute and a half at best**, and unbounded at
worst, before any of the mesh's own work begins.
## The judgement
**This is too slow for an inner loop**, and the reason is not the design.
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) argues that
making the bootstrap path the inner development loop turns the least-exercised code in the
system into the most-exercised. That argument holds only while resetting is cheap. At a minute
and a half a cycle, with occasional multi-minute stalls, the loop is one a person works around
— and the path stays under-exercised for exactly the reason it always was.
Nothing about virtual machines causes this. Hardware virtualisation is present and the machines
boot in fourteen seconds. **The cost is entirely the storage driver**, and the driver is absent
because one userspace package is not installed on the host.
The mesh already has the mechanism for that: a module declares a package, and a hook makes it a
working capability — which is precisely what was just done for the virtualisation daemon
itself, and what [`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)
is about.
## The fix, measured
The comparison was subsequently run. One package — `btrfs-progs`, no dependencies — installed by
hand, the daemon restarted so it re-detected drivers, a copy-on-write pool created on a loop
file, and the identical image launched onto it.
| Operation | `dir` | copy-on-write | |
|---|---|---|---|
| snapshot | 9.9 s, then **> 120 s** | **0.13 s** | ~76× faster, and *consistent* |
| snapshot again | — | 0.12 s | |
| snapshot a third time | — | 0.13 s | |
| restore call | 10.4 s | **0.80 s** | ~13× faster |
| restore, to usable | 20.1 s | **10.5 s** | the remainder is boot, which is irreducible |
| three snapshots, storage | ~4.8 GB | **1.36 GB total, shared** | cost is the delta, not the disk |
**The fix is real, and larger than expected.** Snapshot goes from ten seconds to a tenth of a
second, and — more importantly — from *wildly variable* to *flat*. Three consecutive snapshots
took 0.13, 0.12 and 0.13 seconds. On `dir` the second snapshot never finished.
Storage stops scaling with the machine and starts scaling with what changed: three snapshots of
a 1.5 GB instance occupied 1.36 GB in total, because they share.
### What it projects to
A four-machine reset-and-rerun cycle, the operation the inner loop repeats most:
| | `dir` | copy-on-write |
|---|---|---|
| snapshot the scenario | ~40 s, 6.4 GB | **~0.5 s**, delta-sized |
| restore it | ~40 s + boot | **~3 s** + boot |
| **cycle** | **~90 s, unbounded at worst** | **~15 s, dominated by boot** |
At fifteen seconds, dominated by a boot that cannot be avoided, the inner loop is viable and
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)'s argument holds.
At ninety it did not.
### One honest counter-observation
Launching onto the fresh copy-on-write pool was **slower** — 20.2 s to usable against 14.3 s —
because the image had to be unpacked into a pool that had never seen it. That cost is paid once
per pool, not per scenario, and it is dwarfed by what snapshotting saves. But it is a real
number and it went the other way.
### State this left behind
Recorded because hand-made state is exactly what the mesh's rules exist to prevent, and it must
be declared properly rather than left as an artefact of a measurement:
- `btrfs-progs` installed by hand. Its installation regenerated the boot initramfs, a side
effect worth knowing about.
- The daemon restarted once, to re-detect drivers.
- The test pool and instance were **removed**; the pool the lab actually needs does not exist.