The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18, 19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only the archaeology of what used to be there. Renumbered contiguously. Renames run in ascending order, so every target number is already free and no two files ever collide. The reference rewrite is one simultaneous pass rather than a sequence of replacements. Numbers moved into slots other numbers were vacating -- the node host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time would have cascaded and silently pointed things at the wrong record. Seven plain-text references survived the merges as prose rather than links, naming records that no longer existed: the enrolment token, the link boundary, what a declaration is, reachability, the repository structure. Each mapped to the consolidated record that now holds it. Verified rather than assumed: every [ADR NNNN](path) link now has matching text and target, checked across the whole repository, and the checker passes. Frontmatter `consolidates:` lists dropped -- they named records that are gone, and each consolidated record already says in prose what it absorbed.
61 lines
3.4 KiB
Markdown
61 lines
3.4 KiB
Markdown
---
|
|
status: active
|
|
initiated: 2026-08-24
|
|
touches:
|
|
- 03-DESIGN/01-to-be/03-scenario-lifecycle.md
|
|
- 02-DECISIONS/0009-the-lab.md
|
|
- 02-DECISIONS/0009-the-lab.md
|
|
became: []
|
|
---
|
|
|
|
# 010 — What the lab's inner loop actually costs
|
|
|
|
## What is being investigated
|
|
|
|
[The lifecycle design](../../03-DESIGN/01-to-be/03-scenario-lifecycle.md) closes on an open
|
|
question that is measurable rather than arguable:
|
|
|
|
> **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the
|
|
> inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop
|
|
> is slow, and everything above is theory.
|
|
|
|
Measured on a workstation, 2026-08-24. Numbers in [`measurements.md`](measurements.md).
|
|
|
|
## Why it matters
|
|
|
|
[ADR 0009](../../02-DECISIONS/0009-the-lab.md) makes the
|
|
bootstrap scenario the inner development loop for tiers 0 and 1 — the argument being that
|
|
raising a node from nothing stops being the least-exercised path and becomes the most-exercised
|
|
one. **That argument is only true if raising and resetting are cheap.** A loop that costs
|
|
minutes is a loop people avoid, and the least-exercised path stays least-exercised.
|
|
|
|
## Status
|
|
|
|
Measured, and the finding is a blocker rather than a data point.
|
|
|
|
**A snapshot on the current host is a full copy of the machine's disk.** 1.6 GB and ten seconds
|
|
for one small virtual machine at best — and over two minutes when observed a second time. Cost
|
|
scales with the number of machines and the size of their disks, not with what changed.
|
|
|
|
**With copy-on-write it is 0.13 seconds and costs the delta.** Verified, not assumed. The
|
|
projected four-machine reset cycle falls from roughly ninety seconds to roughly fifteen, of
|
|
which almost all is a boot that cannot be avoided.
|
|
|
|
The cause is not virtual machines and not incus. It is that the host offers incus exactly one
|
|
storage driver, `dir`, which has no copy-on-write and therefore no cheap snapshot. The kernel
|
|
supports btrfs; the userspace tool that would let incus use it is simply not installed.
|
|
|
|
So the question *"is the lab's inner loop fast enough"* currently answers itself the wrong way,
|
|
for a reason that is one declared package away from being fixed — and declaring packages is
|
|
something the mesh already does.
|
|
|
|
## Open questions
|
|
|
|
| Question | Why it matters |
|
|
|---|---|
|
|
| ~~How much does a copy-on-write pool actually improve it?~~ **Measured: snapshot 9.9 s → 0.13 s, restore 10.4 s → 0.80 s, three snapshots sharing 1.36 GB rather than costing 4.8 GB.** The projected four-machine cycle falls from ~90 s to ~15 s. | Answered. The inner-loop argument holds *with* copy-on-write and did not without it. |
|
|
| How does the lab install its own prerequisites on a clean machine? | The lab needs a virtualisation daemon, copy-on-write tooling and a pool before it can do anything — and it cannot depend on the mesh for them, since it is where the mesh is built. |
|
|
| Why was the second snapshot more than twelve times slower than the first? | If snapshot cost is unpredictable rather than merely high, that is worse — a loop with a variable multi-minute step is one nobody trusts. |
|
|
| Does a scenario snapshot need the machines stopped? | Stateless snapshots of a running virtual machine capture the disk but not memory. Whether a mesh restored that way is coherent is not established. |
|
|
| What is the cost at scenario scale — four machines rather than one? | Only single-machine numbers were taken. If the operation is serial, four machines is four times the wait. |
|