The scenario model, the lifecycle, and what the lab actually costs #6
@@ -0,0 +1,55 @@
|
||||
---
|
||||
status: active
|
||||
initiated: 2026-08-24
|
||||
touches:
|
||||
- 03-DESIGN/01-to-be/03-scenario-lifecycle.md
|
||||
- 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
|
||||
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
|
||||
became: []
|
||||
---
|
||||
|
||||
# 010 — What the lab's inner loop actually costs
|
||||
|
||||
## What is being investigated
|
||||
|
||||
[The lifecycle design](../../03-DESIGN/01-to-be/03-scenario-lifecycle.md) closes on an open
|
||||
question that is measurable rather than arguable:
|
||||
|
||||
> **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the
|
||||
> inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop
|
||||
> is slow, and everything above is theory.
|
||||
|
||||
Measured on a workstation, 2026-08-24. Numbers in [`measurements.md`](measurements.md).
|
||||
|
||||
## Why it matters
|
||||
|
||||
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) makes the
|
||||
bootstrap scenario the inner development loop for tiers 0 and 1 — the argument being that
|
||||
raising a node from nothing stops being the least-exercised path and becomes the most-exercised
|
||||
one. **That argument is only true if raising and resetting are cheap.** A loop that costs
|
||||
minutes is a loop people avoid, and the least-exercised path stays least-exercised.
|
||||
|
||||
## Status
|
||||
|
||||
Measured, and the finding is a blocker rather than a data point.
|
||||
|
||||
**A snapshot on the current host is a full copy of the machine's disk.** 1.6 GB and ten seconds
|
||||
for one small virtual machine at best — and over two minutes when observed a second time. Cost
|
||||
scales with the number of machines and the size of their disks, not with what changed.
|
||||
|
||||
The cause is not virtual machines and not incus. It is that the host offers incus exactly one
|
||||
storage driver, `dir`, which has no copy-on-write and therefore no cheap snapshot. The kernel
|
||||
supports btrfs; the userspace tool that would let incus use it is simply not installed.
|
||||
|
||||
So the question *"is the lab's inner loop fast enough"* currently answers itself the wrong way,
|
||||
for a reason that is one declared package away from being fixed — and declaring packages is
|
||||
something the mesh already does.
|
||||
|
||||
## Open questions
|
||||
|
||||
| Question | Why it matters |
|
||||
|---|---|
|
||||
| How much does a copy-on-write pool actually improve it? Expected to be near-instant snapshots and delta-sized storage, but **expected is not measured**. | The whole inner-loop argument rests on the answer. |
|
||||
| Why was the second snapshot more than twelve times slower than the first? | If snapshot cost is unpredictable rather than merely high, that is worse — a loop with a variable multi-minute step is one nobody trusts. |
|
||||
| Does a scenario snapshot need the machines stopped? | Stateless snapshots of a running virtual machine capture the disk but not memory. Whether a mesh restored that way is coherent is not established. |
|
||||
| What is the cost at scenario scale — four machines rather than one? | Only single-machine numbers were taken. If the operation is serial, four machines is four times the wait. |
|
||||
@@ -0,0 +1,98 @@
|
||||
---
|
||||
effort: 010-lab-inner-loop-cost
|
||||
updated: 2026-08-24
|
||||
---
|
||||
|
||||
# Measurements
|
||||
|
||||
Taken 2026-08-24 on a workstation with hardware virtualisation available, an NVMe-backed ext4
|
||||
root, and 300 GB free. One virtual machine, 1 GiB memory, 2 CPUs, from a cached distribution
|
||||
image.
|
||||
|
||||
## The environment, before anything ran
|
||||
|
||||
| Fact | Value | Consequence |
|
||||
|---|---|---|
|
||||
| Hardware virtualisation | present | virtual machines run at native speed; the choice in [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md) is not paying an emulation penalty |
|
||||
| Storage drivers the daemon offers | **`dir` only** | no copy-on-write, therefore no cheap snapshot |
|
||||
| Host filesystems | ext4 throughout | nothing copy-on-write to put a pool on |
|
||||
| btrfs kernel module | **available** | the kernel can do it |
|
||||
| `btrfs-progs` | **not installed** | which is the entire reason the driver is absent |
|
||||
|
||||
The last two rows are the finding. The daemon advertises only `dir` because the userspace tool
|
||||
for anything better is missing — not because the host cannot do better.
|
||||
|
||||
## Raising a machine
|
||||
|
||||
| Step | Time |
|
||||
|---|---|
|
||||
| launch call returns | **3.4 s** |
|
||||
| machine actually usable — a command executes on it | **14.3 s** |
|
||||
|
||||
The gap matters for the lifecycle design: `raise` returning is not the same as the scenario
|
||||
being ready, so the verb has to wait for the second number, not report the first. Reporting
|
||||
the first would be the mesh's own recurring failure — transport reported as effect.
|
||||
|
||||
## Snapshot and restore
|
||||
|
||||
| Operation | Time | Disk |
|
||||
|---|---|---|
|
||||
| snapshot, first | **9.9 s** | **+1.6 GB** |
|
||||
| snapshot, second | **> 120 s — did not complete** | — |
|
||||
| restore call returns | **10.4 s** | — |
|
||||
| machine usable again | **20.1 s** total | — |
|
||||
|
||||
Instance on disk before snapshotting: 1.5 GB. Snapshot directory afterwards: 1.6 GB. **A `dir`
|
||||
snapshot is a full copy** — the storage cost equals the instance, and nothing is shared.
|
||||
|
||||
Implied copy throughput on the first snapshot is roughly 160 MB/s, which is far below what the
|
||||
underlying NVMe can do and is consistent with a real, durable copy rather than a metadata
|
||||
operation.
|
||||
|
||||
**The second snapshot is the more troubling number.** It exceeded two minutes and was still
|
||||
running when the observation was cut off; only the first snapshot exists. Whatever the cause —
|
||||
page cache exhausted by the preceding restore, writeback contention — the practical
|
||||
consequence is that snapshot cost here is **not merely high, it is unpredictable**.
|
||||
|
||||
## What this projects to
|
||||
|
||||
A four-machine scenario, taking the optimistic single-machine numbers and assuming the
|
||||
operations are serial:
|
||||
|
||||
| | one machine | four machines |
|
||||
|---|---|---|
|
||||
| raise, to usable | 14 s | ~57 s |
|
||||
| snapshot | 10 s, 1.6 GB | ~40 s, 6.4 GB |
|
||||
| restore, to usable | 20 s | ~80 s |
|
||||
|
||||
A reset-and-rerun cycle is therefore **around a minute and a half at best**, and unbounded at
|
||||
worst, before any of the mesh's own work begins.
|
||||
|
||||
## The judgement
|
||||
|
||||
**This is too slow for an inner loop**, and the reason is not the design.
|
||||
|
||||
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) argues that
|
||||
making the bootstrap path the inner development loop turns the least-exercised code in the
|
||||
system into the most-exercised. That argument holds only while resetting is cheap. At a minute
|
||||
and a half a cycle, with occasional multi-minute stalls, the loop is one a person works around
|
||||
— and the path stays under-exercised for exactly the reason it always was.
|
||||
|
||||
Nothing about virtual machines causes this. Hardware virtualisation is present and the machines
|
||||
boot in fourteen seconds. **The cost is entirely the storage driver**, and the driver is absent
|
||||
because one userspace package is not installed on the host.
|
||||
|
||||
The mesh already has the mechanism for that: a module declares a package, and a hook makes it a
|
||||
working capability — which is precisely what was just done for the virtualisation daemon
|
||||
itself, and what [`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)
|
||||
is about.
|
||||
|
||||
## Not measured, and it matters
|
||||
|
||||
The copy-on-write comparison **was not run**, because running it would mean installing a package
|
||||
by hand, which the mesh's rules forbid and which would have made the measurement unreproducible
|
||||
anyway. Copy-on-write snapshots are expected to be near-instant with storage proportional to
|
||||
what changed. That expectation is well founded and is still an expectation.
|
||||
|
||||
Until it is measured, the correct statement is: *the current configuration is too slow, and the
|
||||
likely fix is known but unverified.*
|
||||
@@ -48,6 +48,10 @@ The order is not arbitrary — each step needs the one before it to exist:
|
||||
5. **Placement.** Artifacts onto machines.
|
||||
6. **Snapshot**, if the declaration named one.
|
||||
|
||||
`raise` must wait for the machines to be **usable**, not for the call to return. Measured, those
|
||||
are 14 seconds apart for a single machine — and reporting the earlier number would be the mesh's
|
||||
own recurring failure, transport reported as effect.
|
||||
|
||||
Raising is **convergent, not incremental**: raising an instance that already exists brings it to
|
||||
the declared state rather than failing or duplicating. That is the same model the mesh itself
|
||||
uses, and a lab that behaved differently from the thing it tests would be teaching the wrong
|
||||
@@ -122,9 +126,11 @@ decision rather than a second implementation.
|
||||
|
||||
## Open
|
||||
|
||||
- **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the
|
||||
inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop
|
||||
is slow, and everything above is theory.
|
||||
- **What a snapshot costs — measured, and the answer is a blocker.** On the current host a
|
||||
snapshot is a full copy: 1.6 GB and ten seconds for one small machine, over two minutes when
|
||||
observed again, projecting to roughly a minute and a half for a four-machine reset-and-rerun
|
||||
cycle. The cause is the storage driver, not virtual machines. See
|
||||
[research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md).
|
||||
- **Instance naming.** A declaration is a kind and instances are many; how they are named
|
||||
decides whether a person can find the one they left standing yesterday.
|
||||
- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying
|
||||
|
||||
Reference in New Issue
Block a user