The scenario model, the lifecycle, and what the lab actually costs #6
@@ -0,0 +1,55 @@
|
|||||||
|
---
|
||||||
|
status: active
|
||||||
|
initiated: 2026-08-24
|
||||||
|
touches:
|
||||||
|
- 03-DESIGN/01-to-be/03-scenario-lifecycle.md
|
||||||
|
- 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
|
||||||
|
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
|
||||||
|
became: []
|
||||||
|
---
|
||||||
|
|
||||||
|
# 010 — What the lab's inner loop actually costs
|
||||||
|
|
||||||
|
## What is being investigated
|
||||||
|
|
||||||
|
[The lifecycle design](../../03-DESIGN/01-to-be/03-scenario-lifecycle.md) closes on an open
|
||||||
|
question that is measurable rather than arguable:
|
||||||
|
|
||||||
|
> **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the
|
||||||
|
> inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop
|
||||||
|
> is slow, and everything above is theory.
|
||||||
|
|
||||||
|
Measured on a workstation, 2026-08-24. Numbers in [`measurements.md`](measurements.md).
|
||||||
|
|
||||||
|
## Why it matters
|
||||||
|
|
||||||
|
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) makes the
|
||||||
|
bootstrap scenario the inner development loop for tiers 0 and 1 — the argument being that
|
||||||
|
raising a node from nothing stops being the least-exercised path and becomes the most-exercised
|
||||||
|
one. **That argument is only true if raising and resetting are cheap.** A loop that costs
|
||||||
|
minutes is a loop people avoid, and the least-exercised path stays least-exercised.
|
||||||
|
|
||||||
|
## Status
|
||||||
|
|
||||||
|
Measured, and the finding is a blocker rather than a data point.
|
||||||
|
|
||||||
|
**A snapshot on the current host is a full copy of the machine's disk.** 1.6 GB and ten seconds
|
||||||
|
for one small virtual machine at best — and over two minutes when observed a second time. Cost
|
||||||
|
scales with the number of machines and the size of their disks, not with what changed.
|
||||||
|
|
||||||
|
The cause is not virtual machines and not incus. It is that the host offers incus exactly one
|
||||||
|
storage driver, `dir`, which has no copy-on-write and therefore no cheap snapshot. The kernel
|
||||||
|
supports btrfs; the userspace tool that would let incus use it is simply not installed.
|
||||||
|
|
||||||
|
So the question *"is the lab's inner loop fast enough"* currently answers itself the wrong way,
|
||||||
|
for a reason that is one declared package away from being fixed — and declaring packages is
|
||||||
|
something the mesh already does.
|
||||||
|
|
||||||
|
## Open questions
|
||||||
|
|
||||||
|
| Question | Why it matters |
|
||||||
|
|---|---|
|
||||||
|
| How much does a copy-on-write pool actually improve it? Expected to be near-instant snapshots and delta-sized storage, but **expected is not measured**. | The whole inner-loop argument rests on the answer. |
|
||||||
|
| Why was the second snapshot more than twelve times slower than the first? | If snapshot cost is unpredictable rather than merely high, that is worse — a loop with a variable multi-minute step is one nobody trusts. |
|
||||||
|
| Does a scenario snapshot need the machines stopped? | Stateless snapshots of a running virtual machine capture the disk but not memory. Whether a mesh restored that way is coherent is not established. |
|
||||||
|
| What is the cost at scenario scale — four machines rather than one? | Only single-machine numbers were taken. If the operation is serial, four machines is four times the wait. |
|
||||||
@@ -0,0 +1,98 @@
|
|||||||
|
---
|
||||||
|
effort: 010-lab-inner-loop-cost
|
||||||
|
updated: 2026-08-24
|
||||||
|
---
|
||||||
|
|
||||||
|
# Measurements
|
||||||
|
|
||||||
|
Taken 2026-08-24 on a workstation with hardware virtualisation available, an NVMe-backed ext4
|
||||||
|
root, and 300 GB free. One virtual machine, 1 GiB memory, 2 CPUs, from a cached distribution
|
||||||
|
image.
|
||||||
|
|
||||||
|
## The environment, before anything ran
|
||||||
|
|
||||||
|
| Fact | Value | Consequence |
|
||||||
|
|---|---|---|
|
||||||
|
| Hardware virtualisation | present | virtual machines run at native speed; the choice in [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md) is not paying an emulation penalty |
|
||||||
|
| Storage drivers the daemon offers | **`dir` only** | no copy-on-write, therefore no cheap snapshot |
|
||||||
|
| Host filesystems | ext4 throughout | nothing copy-on-write to put a pool on |
|
||||||
|
| btrfs kernel module | **available** | the kernel can do it |
|
||||||
|
| `btrfs-progs` | **not installed** | which is the entire reason the driver is absent |
|
||||||
|
|
||||||
|
The last two rows are the finding. The daemon advertises only `dir` because the userspace tool
|
||||||
|
for anything better is missing — not because the host cannot do better.
|
||||||
|
|
||||||
|
## Raising a machine
|
||||||
|
|
||||||
|
| Step | Time |
|
||||||
|
|---|---|
|
||||||
|
| launch call returns | **3.4 s** |
|
||||||
|
| machine actually usable — a command executes on it | **14.3 s** |
|
||||||
|
|
||||||
|
The gap matters for the lifecycle design: `raise` returning is not the same as the scenario
|
||||||
|
being ready, so the verb has to wait for the second number, not report the first. Reporting
|
||||||
|
the first would be the mesh's own recurring failure — transport reported as effect.
|
||||||
|
|
||||||
|
## Snapshot and restore
|
||||||
|
|
||||||
|
| Operation | Time | Disk |
|
||||||
|
|---|---|---|
|
||||||
|
| snapshot, first | **9.9 s** | **+1.6 GB** |
|
||||||
|
| snapshot, second | **> 120 s — did not complete** | — |
|
||||||
|
| restore call returns | **10.4 s** | — |
|
||||||
|
| machine usable again | **20.1 s** total | — |
|
||||||
|
|
||||||
|
Instance on disk before snapshotting: 1.5 GB. Snapshot directory afterwards: 1.6 GB. **A `dir`
|
||||||
|
snapshot is a full copy** — the storage cost equals the instance, and nothing is shared.
|
||||||
|
|
||||||
|
Implied copy throughput on the first snapshot is roughly 160 MB/s, which is far below what the
|
||||||
|
underlying NVMe can do and is consistent with a real, durable copy rather than a metadata
|
||||||
|
operation.
|
||||||
|
|
||||||
|
**The second snapshot is the more troubling number.** It exceeded two minutes and was still
|
||||||
|
running when the observation was cut off; only the first snapshot exists. Whatever the cause —
|
||||||
|
page cache exhausted by the preceding restore, writeback contention — the practical
|
||||||
|
consequence is that snapshot cost here is **not merely high, it is unpredictable**.
|
||||||
|
|
||||||
|
## What this projects to
|
||||||
|
|
||||||
|
A four-machine scenario, taking the optimistic single-machine numbers and assuming the
|
||||||
|
operations are serial:
|
||||||
|
|
||||||
|
| | one machine | four machines |
|
||||||
|
|---|---|---|
|
||||||
|
| raise, to usable | 14 s | ~57 s |
|
||||||
|
| snapshot | 10 s, 1.6 GB | ~40 s, 6.4 GB |
|
||||||
|
| restore, to usable | 20 s | ~80 s |
|
||||||
|
|
||||||
|
A reset-and-rerun cycle is therefore **around a minute and a half at best**, and unbounded at
|
||||||
|
worst, before any of the mesh's own work begins.
|
||||||
|
|
||||||
|
## The judgement
|
||||||
|
|
||||||
|
**This is too slow for an inner loop**, and the reason is not the design.
|
||||||
|
|
||||||
|
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) argues that
|
||||||
|
making the bootstrap path the inner development loop turns the least-exercised code in the
|
||||||
|
system into the most-exercised. That argument holds only while resetting is cheap. At a minute
|
||||||
|
and a half a cycle, with occasional multi-minute stalls, the loop is one a person works around
|
||||||
|
— and the path stays under-exercised for exactly the reason it always was.
|
||||||
|
|
||||||
|
Nothing about virtual machines causes this. Hardware virtualisation is present and the machines
|
||||||
|
boot in fourteen seconds. **The cost is entirely the storage driver**, and the driver is absent
|
||||||
|
because one userspace package is not installed on the host.
|
||||||
|
|
||||||
|
The mesh already has the mechanism for that: a module declares a package, and a hook makes it a
|
||||||
|
working capability — which is precisely what was just done for the virtualisation daemon
|
||||||
|
itself, and what [`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)
|
||||||
|
is about.
|
||||||
|
|
||||||
|
## Not measured, and it matters
|
||||||
|
|
||||||
|
The copy-on-write comparison **was not run**, because running it would mean installing a package
|
||||||
|
by hand, which the mesh's rules forbid and which would have made the measurement unreproducible
|
||||||
|
anyway. Copy-on-write snapshots are expected to be near-instant with storage proportional to
|
||||||
|
what changed. That expectation is well founded and is still an expectation.
|
||||||
|
|
||||||
|
Until it is measured, the correct statement is: *the current configuration is too slow, and the
|
||||||
|
likely fix is known but unverified.*
|
||||||
@@ -48,6 +48,10 @@ The order is not arbitrary — each step needs the one before it to exist:
|
|||||||
5. **Placement.** Artifacts onto machines.
|
5. **Placement.** Artifacts onto machines.
|
||||||
6. **Snapshot**, if the declaration named one.
|
6. **Snapshot**, if the declaration named one.
|
||||||
|
|
||||||
|
`raise` must wait for the machines to be **usable**, not for the call to return. Measured, those
|
||||||
|
are 14 seconds apart for a single machine — and reporting the earlier number would be the mesh's
|
||||||
|
own recurring failure, transport reported as effect.
|
||||||
|
|
||||||
Raising is **convergent, not incremental**: raising an instance that already exists brings it to
|
Raising is **convergent, not incremental**: raising an instance that already exists brings it to
|
||||||
the declared state rather than failing or duplicating. That is the same model the mesh itself
|
the declared state rather than failing or duplicating. That is the same model the mesh itself
|
||||||
uses, and a lab that behaved differently from the thing it tests would be teaching the wrong
|
uses, and a lab that behaved differently from the thing it tests would be teaching the wrong
|
||||||
@@ -122,9 +126,11 @@ decision rather than a second implementation.
|
|||||||
|
|
||||||
## Open
|
## Open
|
||||||
|
|
||||||
- **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the
|
- **What a snapshot costs — measured, and the answer is a blocker.** On the current host a
|
||||||
inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop
|
snapshot is a full copy: 1.6 GB and ten seconds for one small machine, over two minutes when
|
||||||
is slow, and everything above is theory.
|
observed again, projecting to roughly a minute and a half for a four-machine reset-and-rerun
|
||||||
|
cycle. The cause is the storage driver, not virtual machines. See
|
||||||
|
[research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md).
|
||||||
- **Instance naming.** A declaration is a kind and instances are many; how they are named
|
- **Instance naming.** A declaration is a kind and instances are many; how they are named
|
||||||
decides whether a person can find the one they left standing yesterday.
|
decides whether a person can find the one they left standing yesterday.
|
||||||
- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying
|
- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying
|
||||||
|
|||||||
Reference in New Issue
Block a user