Measure the lab's inner loop — it is too slow, for a fixable reason

The lifecycle design closed on a question that was measurable rather than
arguable, so it was measured. One virtual machine on a workstation with
hardware virtualisation and NVMe.

Raising: the launch call returns in 3.4s, the machine is actually usable
after 14.3s. The gap is a design constraint — raise must wait for the
second number, because reporting the first would be transport reported as
effect, which is the mesh's own recurring failure.

Snapshot: 9.9s and 1.6 GB for a 1.5 GB instance. A dir snapshot is a full
copy; nothing is shared. Restore: 10.4s, usable again after 20.1s.

The second snapshot exceeded two minutes and never completed. That is the
more troubling number: snapshot cost here is not merely high, it is
unpredictable, and a loop with a variable multi-minute step is one nobody
trusts.

Projected to a four-machine scenario, a reset-and-rerun cycle is about a
minute and a half at best and unbounded at worst, before any of the mesh's
own work begins. That is too slow for an inner loop, and ADR 0029's whole
argument — that making the bootstrap path the inner loop turns the
least-exercised code into the most-exercised — holds only while resetting
is cheap.

The cause is not virtual machines. Hardware virtualisation is present and
machines boot in fourteen seconds. It is that the daemon offers exactly one
storage driver, dir, which has no copy-on-write and therefore no cheap
snapshot. The btrfs kernel module is available; btrfs-progs is simply not
installed, which is the entire reason the driver is absent.

The copy-on-write comparison was deliberately NOT run, because running it
would mean installing a package by hand — which the rules forbid and which
would have made the measurement unreproducible. So the honest statement is
that the current configuration is too slow and the likely fix is known but
unverified, rather than that btrfs fixes it.
This commit is contained in:
2026-08-24 00:07:22 +02:00
parent a253afe020
commit 98bcd5cc49
3 changed files with 162 additions and 3 deletions
+9 -3
View File
@@ -48,6 +48,10 @@ The order is not arbitrary — each step needs the one before it to exist:
5. **Placement.** Artifacts onto machines.
6. **Snapshot**, if the declaration named one.
`raise` must wait for the machines to be **usable**, not for the call to return. Measured, those
are 14 seconds apart for a single machine — and reporting the earlier number would be the mesh's
own recurring failure, transport reported as effect.
Raising is **convergent, not incremental**: raising an instance that already exists brings it to
the declared state rather than failing or duplicating. That is the same model the mesh itself
uses, and a lab that behaved differently from the thing it tests would be teaching the wrong
@@ -122,9 +126,11 @@ decision rather than a second implementation.
## Open
- **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the
inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop
is slow, and everything above is theory.
- **What a snapshot costs — measured, and the answer is a blocker.** On the current host a
snapshot is a full copy: 1.6 GB and ten seconds for one small machine, over two minutes when
observed again, projecting to roughly a minute and a half for a four-machine reset-and-rerun
cycle. The cause is the storage driver, not virtual machines. See
[research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md).
- **Instance naming.** A declaration is a kind and instances are many; how they are named
decides whether a person can find the one they left standing yesterday.
- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying