Files
hq/01-RESEARCH/010-lab-inner-loop-cost/measurements.md
T
jschoubben 98bcd5cc49 Measure the lab's inner loop — it is too slow, for a fixable reason
The lifecycle design closed on a question that was measurable rather than
arguable, so it was measured. One virtual machine on a workstation with
hardware virtualisation and NVMe.

Raising: the launch call returns in 3.4s, the machine is actually usable
after 14.3s. The gap is a design constraint — raise must wait for the
second number, because reporting the first would be transport reported as
effect, which is the mesh's own recurring failure.

Snapshot: 9.9s and 1.6 GB for a 1.5 GB instance. A dir snapshot is a full
copy; nothing is shared. Restore: 10.4s, usable again after 20.1s.

The second snapshot exceeded two minutes and never completed. That is the
more troubling number: snapshot cost here is not merely high, it is
unpredictable, and a loop with a variable multi-minute step is one nobody
trusts.

Projected to a four-machine scenario, a reset-and-rerun cycle is about a
minute and a half at best and unbounded at worst, before any of the mesh's
own work begins. That is too slow for an inner loop, and ADR 0029's whole
argument — that making the bootstrap path the inner loop turns the
least-exercised code into the most-exercised — holds only while resetting
is cheap.

The cause is not virtual machines. Hardware virtualisation is present and
machines boot in fourteen seconds. It is that the daemon offers exactly one
storage driver, dir, which has no copy-on-write and therefore no cheap
snapshot. The btrfs kernel module is available; btrfs-progs is simply not
installed, which is the entire reason the driver is absent.

The copy-on-write comparison was deliberately NOT run, because running it
would mean installing a package by hand — which the rules forbid and which
would have made the measurement unreproducible. So the honest statement is
that the current configuration is too slow and the likely fix is known but
unverified, rather than that btrfs fixes it.
2026-08-24 00:07:22 +02:00

4.5 KiB

effort, updated
effort updated
010-lab-inner-loop-cost 2026-08-24

Measurements

Taken 2026-08-24 on a workstation with hardware virtualisation available, an NVMe-backed ext4 root, and 300 GB free. One virtual machine, 1 GiB memory, 2 CPUs, from a cached distribution image.

The environment, before anything ran

Fact Value Consequence
Hardware virtualisation present virtual machines run at native speed; the choice in ADR 0016 is not paying an emulation penalty
Storage drivers the daemon offers dir only no copy-on-write, therefore no cheap snapshot
Host filesystems ext4 throughout nothing copy-on-write to put a pool on
btrfs kernel module available the kernel can do it
btrfs-progs not installed which is the entire reason the driver is absent

The last two rows are the finding. The daemon advertises only dir because the userspace tool for anything better is missing — not because the host cannot do better.

Raising a machine

Step Time
launch call returns 3.4 s
machine actually usable — a command executes on it 14.3 s

The gap matters for the lifecycle design: raise returning is not the same as the scenario being ready, so the verb has to wait for the second number, not report the first. Reporting the first would be the mesh's own recurring failure — transport reported as effect.

Snapshot and restore

Operation Time Disk
snapshot, first 9.9 s +1.6 GB
snapshot, second > 120 s — did not complete —
restore call returns 10.4 s —
machine usable again 20.1 s total —

Instance on disk before snapshotting: 1.5 GB. Snapshot directory afterwards: 1.6 GB. A dir snapshot is a full copy — the storage cost equals the instance, and nothing is shared.

Implied copy throughput on the first snapshot is roughly 160 MB/s, which is far below what the underlying NVMe can do and is consistent with a real, durable copy rather than a metadata operation.

The second snapshot is the more troubling number. It exceeded two minutes and was still running when the observation was cut off; only the first snapshot exists. Whatever the cause — page cache exhausted by the preceding restore, writeback contention — the practical consequence is that snapshot cost here is not merely high, it is unpredictable.

What this projects to

A four-machine scenario, taking the optimistic single-machine numbers and assuming the operations are serial:

one machine four machines
raise, to usable 14 s ~57 s
snapshot 10 s, 1.6 GB ~40 s, 6.4 GB
restore, to usable 20 s ~80 s

A reset-and-rerun cycle is therefore around a minute and a half at best, and unbounded at worst, before any of the mesh's own work begins.

The judgement

This is too slow for an inner loop, and the reason is not the design.

ADR 0029 argues that making the bootstrap path the inner development loop turns the least-exercised code in the system into the most-exercised. That argument holds only while resetting is cheap. At a minute and a half a cycle, with occasional multi-minute stalls, the loop is one a person works around — and the path stays under-exercised for exactly the reason it always was.

Nothing about virtual machines causes this. Hardware virtualisation is present and the machines boot in fourteen seconds. The cost is entirely the storage driver, and the driver is absent because one userspace package is not installed on the host.

The mesh already has the mechanism for that: a module declares a package, and a hook makes it a working capability — which is precisely what was just done for the virtualisation daemon itself, and what 04-ISSUES/007 is about.

Not measured, and it matters

The copy-on-write comparison was not run, because running it would mean installing a package by hand, which the mesh's rules forbid and which would have made the measurement unreproducible anyway. Copy-on-write snapshots are expected to be near-instant with storage proportional to what changed. That expectation is well founded and is still an expectation.

Until it is measured, the correct statement is: the current configuration is too slow, and the likely fix is known but unverified.