diff --git a/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md b/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md new file mode 100644 index 0000000..bd52bde --- /dev/null +++ b/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md @@ -0,0 +1,55 @@ +--- +status: active +initiated: 2026-08-24 +touches: + - 03-DESIGN/01-to-be/03-scenario-lifecycle.md + - 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md + - 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md +became: [] +--- + +# 010 — What the lab's inner loop actually costs + +## What is being investigated + +[The lifecycle design](../../03-DESIGN/01-to-be/03-scenario-lifecycle.md) closes on an open +question that is measurable rather than arguable: + +> **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the +> inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop +> is slow, and everything above is theory. + +Measured on a workstation, 2026-08-24. Numbers in [`measurements.md`](measurements.md). + +## Why it matters + +[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) makes the +bootstrap scenario the inner development loop for tiers 0 and 1 — the argument being that +raising a node from nothing stops being the least-exercised path and becomes the most-exercised +one. **That argument is only true if raising and resetting are cheap.** A loop that costs +minutes is a loop people avoid, and the least-exercised path stays least-exercised. + +## Status + +Measured, and the finding is a blocker rather than a data point. + +**A snapshot on the current host is a full copy of the machine's disk.** 1.6 GB and ten seconds +for one small virtual machine at best — and over two minutes when observed a second time. Cost +scales with the number of machines and the size of their disks, not with what changed. + +The cause is not virtual machines and not incus. It is that the host offers incus exactly one +storage driver, `dir`, which has no copy-on-write and therefore no cheap snapshot. The kernel +supports btrfs; the userspace tool that would let incus use it is simply not installed. + +So the question *"is the lab's inner loop fast enough"* currently answers itself the wrong way, +for a reason that is one declared package away from being fixed — and declaring packages is +something the mesh already does. + +## Open questions + +| Question | Why it matters | +|---|---| +| How much does a copy-on-write pool actually improve it? Expected to be near-instant snapshots and delta-sized storage, but **expected is not measured**. | The whole inner-loop argument rests on the answer. | +| Why was the second snapshot more than twelve times slower than the first? | If snapshot cost is unpredictable rather than merely high, that is worse — a loop with a variable multi-minute step is one nobody trusts. | +| Does a scenario snapshot need the machines stopped? | Stateless snapshots of a running virtual machine capture the disk but not memory. Whether a mesh restored that way is coherent is not established. | +| What is the cost at scenario scale — four machines rather than one? | Only single-machine numbers were taken. If the operation is serial, four machines is four times the wait. | diff --git a/01-RESEARCH/010-lab-inner-loop-cost/measurements.md b/01-RESEARCH/010-lab-inner-loop-cost/measurements.md new file mode 100644 index 0000000..9189c07 --- /dev/null +++ b/01-RESEARCH/010-lab-inner-loop-cost/measurements.md @@ -0,0 +1,98 @@ +--- +effort: 010-lab-inner-loop-cost +updated: 2026-08-24 +--- + +# Measurements + +Taken 2026-08-24 on a workstation with hardware virtualisation available, an NVMe-backed ext4 +root, and 300 GB free. One virtual machine, 1 GiB memory, 2 CPUs, from a cached distribution +image. + +## The environment, before anything ran + +| Fact | Value | Consequence | +|---|---|---| +| Hardware virtualisation | present | virtual machines run at native speed; the choice in [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md) is not paying an emulation penalty | +| Storage drivers the daemon offers | **`dir` only** | no copy-on-write, therefore no cheap snapshot | +| Host filesystems | ext4 throughout | nothing copy-on-write to put a pool on | +| btrfs kernel module | **available** | the kernel can do it | +| `btrfs-progs` | **not installed** | which is the entire reason the driver is absent | + +The last two rows are the finding. The daemon advertises only `dir` because the userspace tool +for anything better is missing — not because the host cannot do better. + +## Raising a machine + +| Step | Time | +|---|---| +| launch call returns | **3.4 s** | +| machine actually usable — a command executes on it | **14.3 s** | + +The gap matters for the lifecycle design: `raise` returning is not the same as the scenario +being ready, so the verb has to wait for the second number, not report the first. Reporting +the first would be the mesh's own recurring failure — transport reported as effect. + +## Snapshot and restore + +| Operation | Time | Disk | +|---|---|---| +| snapshot, first | **9.9 s** | **+1.6 GB** | +| snapshot, second | **> 120 s — did not complete** | — | +| restore call returns | **10.4 s** | — | +| machine usable again | **20.1 s** total | — | + +Instance on disk before snapshotting: 1.5 GB. Snapshot directory afterwards: 1.6 GB. **A `dir` +snapshot is a full copy** — the storage cost equals the instance, and nothing is shared. + +Implied copy throughput on the first snapshot is roughly 160 MB/s, which is far below what the +underlying NVMe can do and is consistent with a real, durable copy rather than a metadata +operation. + +**The second snapshot is the more troubling number.** It exceeded two minutes and was still +running when the observation was cut off; only the first snapshot exists. Whatever the cause — +page cache exhausted by the preceding restore, writeback contention — the practical +consequence is that snapshot cost here is **not merely high, it is unpredictable**. + +## What this projects to + +A four-machine scenario, taking the optimistic single-machine numbers and assuming the +operations are serial: + +| | one machine | four machines | +|---|---|---| +| raise, to usable | 14 s | ~57 s | +| snapshot | 10 s, 1.6 GB | ~40 s, 6.4 GB | +| restore, to usable | 20 s | ~80 s | + +A reset-and-rerun cycle is therefore **around a minute and a half at best**, and unbounded at +worst, before any of the mesh's own work begins. + +## The judgement + +**This is too slow for an inner loop**, and the reason is not the design. + +[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) argues that +making the bootstrap path the inner development loop turns the least-exercised code in the +system into the most-exercised. That argument holds only while resetting is cheap. At a minute +and a half a cycle, with occasional multi-minute stalls, the loop is one a person works around +— and the path stays under-exercised for exactly the reason it always was. + +Nothing about virtual machines causes this. Hardware virtualisation is present and the machines +boot in fourteen seconds. **The cost is entirely the storage driver**, and the driver is absent +because one userspace package is not installed on the host. + +The mesh already has the mechanism for that: a module declares a package, and a hook makes it a +working capability — which is precisely what was just done for the virtualisation daemon +itself, and what [`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md) +is about. + +## Not measured, and it matters + +The copy-on-write comparison **was not run**, because running it would mean installing a package +by hand, which the mesh's rules forbid and which would have made the measurement unreproducible +anyway. Copy-on-write snapshots are expected to be near-instant with storage proportional to +what changed. That expectation is well founded and is still an expectation. + +Until it is measured, the correct statement is: *the current configuration is too slow, and the +likely fix is known but unverified.* diff --git a/03-DESIGN/01-to-be/03-scenario-lifecycle.md b/03-DESIGN/01-to-be/03-scenario-lifecycle.md index 3ca2623..8a06df0 100644 --- a/03-DESIGN/01-to-be/03-scenario-lifecycle.md +++ b/03-DESIGN/01-to-be/03-scenario-lifecycle.md @@ -48,6 +48,10 @@ The order is not arbitrary — each step needs the one before it to exist: 5. **Placement.** Artifacts onto machines. 6. **Snapshot**, if the declaration named one. +`raise` must wait for the machines to be **usable**, not for the call to return. Measured, those +are 14 seconds apart for a single machine — and reporting the earlier number would be the mesh's +own recurring failure, transport reported as effect. + Raising is **convergent, not incremental**: raising an instance that already exists brings it to the declared state rather than failing or duplicating. That is the same model the mesh itself uses, and a lab that behaved differently from the thing it tests would be teaching the wrong @@ -122,9 +126,11 @@ decision rather than a second implementation. ## Open -- **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the - inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop - is slow, and everything above is theory. +- **What a snapshot costs — measured, and the answer is a blocker.** On the current host a + snapshot is a full copy: 1.6 GB and ten seconds for one small machine, over two minutes when + observed again, projecting to roughly a minute and a half for a four-machine reset-and-rerun + cycle. The cause is the storage driver, not virtual machines. See + [research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md). - **Instance naming.** A declaration is a kind and instances are many; how they are named decides whether a person can find the one they left standing yesterday. - **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying