--- status: active initiated: 2026-08-24 touches: - 03-DESIGN/01-to-be/03-scenario-lifecycle.md - 02-DECISIONS/0016-the-lab.md - 02-DECISIONS/0016-the-lab.md became: [] --- # 010 — What the lab's inner loop actually costs ## What is being investigated [The lifecycle design](../../03-DESIGN/01-to-be/03-scenario-lifecycle.md) closes on an open question that is measurable rather than arguable: > **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the > inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop > is slow, and everything above is theory. Measured on a workstation, 2026-08-24. Numbers in [`measurements.md`](measurements.md). ## Why it matters [ADR 0016](../../02-DECISIONS/0016-the-lab.md) makes the bootstrap scenario the inner development loop for tiers 0 and 1 — the argument being that raising a node from nothing stops being the least-exercised path and becomes the most-exercised one. **That argument is only true if raising and resetting are cheap.** A loop that costs minutes is a loop people avoid, and the least-exercised path stays least-exercised. ## Status Measured, and the finding is a blocker rather than a data point. **A snapshot on the current host is a full copy of the machine's disk.** 1.6 GB and ten seconds for one small virtual machine at best — and over two minutes when observed a second time. Cost scales with the number of machines and the size of their disks, not with what changed. **With copy-on-write it is 0.13 seconds and costs the delta.** Verified, not assumed. The projected four-machine reset cycle falls from roughly ninety seconds to roughly fifteen, of which almost all is a boot that cannot be avoided. The cause is not virtual machines and not incus. It is that the host offers incus exactly one storage driver, `dir`, which has no copy-on-write and therefore no cheap snapshot. The kernel supports btrfs; the userspace tool that would let incus use it is simply not installed. So the question *"is the lab's inner loop fast enough"* currently answers itself the wrong way, for a reason that is one declared package away from being fixed — and declaring packages is something the mesh already does. ## Open questions | Question | Why it matters | |---|---| | ~~How much does a copy-on-write pool actually improve it?~~ **Measured: snapshot 9.9 s → 0.13 s, restore 10.4 s → 0.80 s, three snapshots sharing 1.36 GB rather than costing 4.8 GB.** The projected four-machine cycle falls from ~90 s to ~15 s. | Answered. The inner-loop argument holds *with* copy-on-write and did not without it. | | How does the lab install its own prerequisites on a clean machine? | The lab needs a virtualisation daemon, copy-on-write tooling and a pool before it can do anything — and it cannot depend on the mesh for them, since it is where the mesh is built. | | Why was the second snapshot more than twelve times slower than the first? | If snapshot cost is unpredictable rather than merely high, that is worse — a loop with a variable multi-minute step is one nobody trusts. | | Does a scenario snapshot need the machines stopped? | Stateless snapshots of a running virtual machine capture the disk but not memory. Whether a mesh restored that way is coherent is not established. | | What is the cost at scenario scale — four machines rather than one? | Only single-machine numbers were taken. If the operation is serial, four machines is four times the wait. |