The lifecycle design closed on a question that was measurable rather than arguable, so it was measured. One virtual machine on a workstation with hardware virtualisation and NVMe. Raising: the launch call returns in 3.4s, the machine is actually usable after 14.3s. The gap is a design constraint — raise must wait for the second number, because reporting the first would be transport reported as effect, which is the mesh's own recurring failure. Snapshot: 9.9s and 1.6 GB for a 1.5 GB instance. A dir snapshot is a full copy; nothing is shared. Restore: 10.4s, usable again after 20.1s. The second snapshot exceeded two minutes and never completed. That is the more troubling number: snapshot cost here is not merely high, it is unpredictable, and a loop with a variable multi-minute step is one nobody trusts. Projected to a four-machine scenario, a reset-and-rerun cycle is about a minute and a half at best and unbounded at worst, before any of the mesh's own work begins. That is too slow for an inner loop, and ADR 0029's whole argument — that making the bootstrap path the inner loop turns the least-exercised code into the most-exercised — holds only while resetting is cheap. The cause is not virtual machines. Hardware virtualisation is present and machines boot in fourteen seconds. It is that the daemon offers exactly one storage driver, dir, which has no copy-on-write and therefore no cheap snapshot. The btrfs kernel module is available; btrfs-progs is simply not installed, which is the entire reason the driver is absent. The copy-on-write comparison was deliberately NOT run, because running it would mean installing a package by hand — which the rules forbid and which would have made the measurement unreproducible. So the honest statement is that the current configuration is too slow and the likely fix is known but unverified, rather than that btrfs fixes it.
56 lines
2.9 KiB
Markdown
56 lines
2.9 KiB
Markdown
---
|
|
status: active
|
|
initiated: 2026-08-24
|
|
touches:
|
|
- 03-DESIGN/01-to-be/03-scenario-lifecycle.md
|
|
- 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
|
|
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
|
|
became: []
|
|
---
|
|
|
|
# 010 — What the lab's inner loop actually costs
|
|
|
|
## What is being investigated
|
|
|
|
[The lifecycle design](../../03-DESIGN/01-to-be/03-scenario-lifecycle.md) closes on an open
|
|
question that is measurable rather than arguable:
|
|
|
|
> **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the
|
|
> inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop
|
|
> is slow, and everything above is theory.
|
|
|
|
Measured on a workstation, 2026-08-24. Numbers in [`measurements.md`](measurements.md).
|
|
|
|
## Why it matters
|
|
|
|
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) makes the
|
|
bootstrap scenario the inner development loop for tiers 0 and 1 — the argument being that
|
|
raising a node from nothing stops being the least-exercised path and becomes the most-exercised
|
|
one. **That argument is only true if raising and resetting are cheap.** A loop that costs
|
|
minutes is a loop people avoid, and the least-exercised path stays least-exercised.
|
|
|
|
## Status
|
|
|
|
Measured, and the finding is a blocker rather than a data point.
|
|
|
|
**A snapshot on the current host is a full copy of the machine's disk.** 1.6 GB and ten seconds
|
|
for one small virtual machine at best — and over two minutes when observed a second time. Cost
|
|
scales with the number of machines and the size of their disks, not with what changed.
|
|
|
|
The cause is not virtual machines and not incus. It is that the host offers incus exactly one
|
|
storage driver, `dir`, which has no copy-on-write and therefore no cheap snapshot. The kernel
|
|
supports btrfs; the userspace tool that would let incus use it is simply not installed.
|
|
|
|
So the question *"is the lab's inner loop fast enough"* currently answers itself the wrong way,
|
|
for a reason that is one declared package away from being fixed — and declaring packages is
|
|
something the mesh already does.
|
|
|
|
## Open questions
|
|
|
|
| Question | Why it matters |
|
|
|---|---|
|
|
| How much does a copy-on-write pool actually improve it? Expected to be near-instant snapshots and delta-sized storage, but **expected is not measured**. | The whole inner-loop argument rests on the answer. |
|
|
| Why was the second snapshot more than twelve times slower than the first? | If snapshot cost is unpredictable rather than merely high, that is worse — a loop with a variable multi-minute step is one nobody trusts. |
|
|
| Does a scenario snapshot need the machines stopped? | Stateless snapshots of a running virtual machine capture the disk but not memory. Whether a mesh restored that way is coherent is not established. |
|
|
| What is the cost at scenario scale — four machines rather than one? | Only single-machine numbers were taken. If the operation is serial, four machines is four times the wait. |
|