Measure the lab's inner loop — it is too slow, for a fixable reason

The lifecycle design closed on a question that was measurable rather than
arguable, so it was measured. One virtual machine on a workstation with
hardware virtualisation and NVMe.

Raising: the launch call returns in 3.4s, the machine is actually usable
after 14.3s. The gap is a design constraint — raise must wait for the
second number, because reporting the first would be transport reported as
effect, which is the mesh's own recurring failure.

Snapshot: 9.9s and 1.6 GB for a 1.5 GB instance. A dir snapshot is a full
copy; nothing is shared. Restore: 10.4s, usable again after 20.1s.

The second snapshot exceeded two minutes and never completed. That is the
more troubling number: snapshot cost here is not merely high, it is
unpredictable, and a loop with a variable multi-minute step is one nobody
trusts.

Projected to a four-machine scenario, a reset-and-rerun cycle is about a
minute and a half at best and unbounded at worst, before any of the mesh's
own work begins. That is too slow for an inner loop, and ADR 0029's whole
argument — that making the bootstrap path the inner loop turns the
least-exercised code into the most-exercised — holds only while resetting
is cheap.

The cause is not virtual machines. Hardware virtualisation is present and
machines boot in fourteen seconds. It is that the daemon offers exactly one
storage driver, dir, which has no copy-on-write and therefore no cheap
snapshot. The btrfs kernel module is available; btrfs-progs is simply not
installed, which is the entire reason the driver is absent.

The copy-on-write comparison was deliberately NOT run, because running it
would mean installing a package by hand — which the rules forbid and which
would have made the measurement unreproducible. So the honest statement is
that the current configuration is too slow and the likely fix is known but
unverified, rather than that btrfs fixes it.
This commit is contained in:
2026-08-24 00:07:22 +02:00
parent a253afe020
commit 98bcd5cc49
3 changed files with 162 additions and 3 deletions
@@ -0,0 +1,55 @@
---
status: active
initiated: 2026-08-24
touches:
- 03-DESIGN/01-to-be/03-scenario-lifecycle.md
- 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
became: []
---
# 010 — What the lab's inner loop actually costs
## What is being investigated
[The lifecycle design](../../03-DESIGN/01-to-be/03-scenario-lifecycle.md) closes on an open
question that is measurable rather than arguable:
> **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the
> inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop
> is slow, and everything above is theory.
Measured on a workstation, 2026-08-24. Numbers in [`measurements.md`](measurements.md).
## Why it matters
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) makes the
bootstrap scenario the inner development loop for tiers 0 and 1 — the argument being that
raising a node from nothing stops being the least-exercised path and becomes the most-exercised
one. **That argument is only true if raising and resetting are cheap.** A loop that costs
minutes is a loop people avoid, and the least-exercised path stays least-exercised.
## Status
Measured, and the finding is a blocker rather than a data point.
**A snapshot on the current host is a full copy of the machine's disk.** 1.6 GB and ten seconds
for one small virtual machine at best — and over two minutes when observed a second time. Cost
scales with the number of machines and the size of their disks, not with what changed.
The cause is not virtual machines and not incus. It is that the host offers incus exactly one
storage driver, `dir`, which has no copy-on-write and therefore no cheap snapshot. The kernel
supports btrfs; the userspace tool that would let incus use it is simply not installed.
So the question *"is the lab's inner loop fast enough"* currently answers itself the wrong way,
for a reason that is one declared package away from being fixed — and declaring packages is
something the mesh already does.
## Open questions
| Question | Why it matters |
|---|---|
| How much does a copy-on-write pool actually improve it? Expected to be near-instant snapshots and delta-sized storage, but **expected is not measured**. | The whole inner-loop argument rests on the answer. |
| Why was the second snapshot more than twelve times slower than the first? | If snapshot cost is unpredictable rather than merely high, that is worse — a loop with a variable multi-minute step is one nobody trusts. |
| Does a scenario snapshot need the machines stopped? | Stateless snapshots of a running virtual machine capture the disk but not memory. Whether a mesh restored that way is coherent is not established. |
| What is the cost at scenario scale — four machines rather than one? | Only single-machine numbers were taken. If the operation is serial, four machines is four times the wait. |
@@ -0,0 +1,98 @@
---
effort: 010-lab-inner-loop-cost
updated: 2026-08-24
---
# Measurements
Taken 2026-08-24 on a workstation with hardware virtualisation available, an NVMe-backed ext4
root, and 300 GB free. One virtual machine, 1 GiB memory, 2 CPUs, from a cached distribution
image.
## The environment, before anything ran
| Fact | Value | Consequence |
|---|---|---|
| Hardware virtualisation | present | virtual machines run at native speed; the choice in [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md) is not paying an emulation penalty |
| Storage drivers the daemon offers | **`dir` only** | no copy-on-write, therefore no cheap snapshot |
| Host filesystems | ext4 throughout | nothing copy-on-write to put a pool on |
| btrfs kernel module | **available** | the kernel can do it |
| `btrfs-progs` | **not installed** | which is the entire reason the driver is absent |
The last two rows are the finding. The daemon advertises only `dir` because the userspace tool
for anything better is missing — not because the host cannot do better.
## Raising a machine
| Step | Time |
|---|---|
| launch call returns | **3.4 s** |
| machine actually usable — a command executes on it | **14.3 s** |
The gap matters for the lifecycle design: `raise` returning is not the same as the scenario
being ready, so the verb has to wait for the second number, not report the first. Reporting
the first would be the mesh's own recurring failure — transport reported as effect.
## Snapshot and restore
| Operation | Time | Disk |
|---|---|---|
| snapshot, first | **9.9 s** | **+1.6 GB** |
| snapshot, second | **> 120 s — did not complete** | — |
| restore call returns | **10.4 s** | — |
| machine usable again | **20.1 s** total | — |
Instance on disk before snapshotting: 1.5 GB. Snapshot directory afterwards: 1.6 GB. **A `dir`
snapshot is a full copy** — the storage cost equals the instance, and nothing is shared.
Implied copy throughput on the first snapshot is roughly 160 MB/s, which is far below what the
underlying NVMe can do and is consistent with a real, durable copy rather than a metadata
operation.
**The second snapshot is the more troubling number.** It exceeded two minutes and was still
running when the observation was cut off; only the first snapshot exists. Whatever the cause —
page cache exhausted by the preceding restore, writeback contention — the practical
consequence is that snapshot cost here is **not merely high, it is unpredictable**.
## What this projects to
A four-machine scenario, taking the optimistic single-machine numbers and assuming the
operations are serial:
| | one machine | four machines |
|---|---|---|
| raise, to usable | 14 s | ~57 s |
| snapshot | 10 s, 1.6 GB | ~40 s, 6.4 GB |
| restore, to usable | 20 s | ~80 s |
A reset-and-rerun cycle is therefore **around a minute and a half at best**, and unbounded at
worst, before any of the mesh's own work begins.
## The judgement
**This is too slow for an inner loop**, and the reason is not the design.
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) argues that
making the bootstrap path the inner development loop turns the least-exercised code in the
system into the most-exercised. That argument holds only while resetting is cheap. At a minute
and a half a cycle, with occasional multi-minute stalls, the loop is one a person works around
— and the path stays under-exercised for exactly the reason it always was.
Nothing about virtual machines causes this. Hardware virtualisation is present and the machines
boot in fourteen seconds. **The cost is entirely the storage driver**, and the driver is absent
because one userspace package is not installed on the host.
The mesh already has the mechanism for that: a module declares a package, and a hook makes it a
working capability — which is precisely what was just done for the virtualisation daemon
itself, and what [`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)
is about.
## Not measured, and it matters
The copy-on-write comparison **was not run**, because running it would mean installing a package
by hand, which the mesh's rules forbid and which would have made the measurement unreproducible
anyway. Copy-on-write snapshots are expected to be near-instant with storage proportional to
what changed. That expectation is well founded and is still an expectation.
Until it is measured, the correct statement is: *the current configuration is too slow, and the
likely fix is known but unverified.*
+9 -3
View File
@@ -48,6 +48,10 @@ The order is not arbitrary — each step needs the one before it to exist:
5. **Placement.** Artifacts onto machines.
6. **Snapshot**, if the declaration named one.
`raise` must wait for the machines to be **usable**, not for the call to return. Measured, those
are 14 seconds apart for a single machine — and reporting the earlier number would be the mesh's
own recurring failure, transport reported as effect.
Raising is **convergent, not incremental**: raising an instance that already exists brings it to
the declared state rather than failing or duplicating. That is the same model the mesh itself
uses, and a lab that behaved differently from the thing it tests would be teaching the wrong
@@ -122,9 +126,11 @@ decision rather than a second implementation.
## Open
- **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the
inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop
is slow, and everything above is theory.
- **What a snapshot costs — measured, and the answer is a blocker.** On the current host a
snapshot is a full copy: 1.6 GB and ten seconds for one small machine, over two minutes when
observed again, projecting to roughly a minute and a half for a four-machine reset-and-rerun
cycle. The cause is the storage driver, not virtual machines. See
[research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md).
- **Instance naming.** A declaration is a kind and instances are many; how they are named
decides whether a person can find the one they left standing yesterday.
- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying