Files
hq/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00

61 lines
3.4 KiB
Markdown

---
status: active
initiated: 2026-08-24
touches:
- 03-DESIGN/01-to-be/03-scenario-lifecycle.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
became: []
---
# 010 — What the lab's inner loop actually costs
## What is being investigated
[The lifecycle design](../../03-DESIGN/01-to-be/03-scenario-lifecycle.md) closes on an open
question that is measurable rather than arguable:
> **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the
> inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop
> is slow, and everything above is theory.
Measured on a workstation, 2026-08-24. Numbers in [`measurements.md`](measurements.md).
## Why it matters
[ADR 0016](../../02-DECISIONS/0016-the-lab.md) makes the
bootstrap scenario the inner development loop for tiers 0 and 1 — the argument being that
raising a node from nothing stops being the least-exercised path and becomes the most-exercised
one. **That argument is only true if raising and resetting are cheap.** A loop that costs
minutes is a loop people avoid, and the least-exercised path stays least-exercised.
## Status
Measured, and the finding is a blocker rather than a data point.
**A snapshot on the current host is a full copy of the machine's disk.** 1.6 GB and ten seconds
for one small virtual machine at best — and over two minutes when observed a second time. Cost
scales with the number of machines and the size of their disks, not with what changed.
**With copy-on-write it is 0.13 seconds and costs the delta.** Verified, not assumed. The
projected four-machine reset cycle falls from roughly ninety seconds to roughly fifteen, of
which almost all is a boot that cannot be avoided.
The cause is not virtual machines and not incus. It is that the host offers incus exactly one
storage driver, `dir`, which has no copy-on-write and therefore no cheap snapshot. The kernel
supports btrfs; the userspace tool that would let incus use it is simply not installed.
So the question *"is the lab's inner loop fast enough"* currently answers itself the wrong way,
for a reason that is one declared package away from being fixed — and declaring packages is
something the mesh already does.
## Open questions
| Question | Why it matters |
|---|---|
| ~~How much does a copy-on-write pool actually improve it?~~ **Measured: snapshot 9.9 s → 0.13 s, restore 10.4 s → 0.80 s, three snapshots sharing 1.36 GB rather than costing 4.8 GB.** The projected four-machine cycle falls from ~90 s to ~15 s. | Answered. The inner-loop argument holds *with* copy-on-write and did not without it. |
| How does the lab install its own prerequisites on a clean machine? | The lab needs a virtualisation daemon, copy-on-write tooling and a pool before it can do anything — and it cannot depend on the mesh for them, since it is where the mesh is built. |
| Why was the second snapshot more than twelve times slower than the first? | If snapshot cost is unpredictable rather than merely high, that is worse — a loop with a variable multi-minute step is one nobody trusts. |
| Does a scenario snapshot need the machines stopped? | Stateless snapshots of a running virtual machine capture the disk but not memory. Whether a mesh restored that way is coherent is not established. |
| What is the cost at scenario scale — four machines rather than one? | Only single-machine numbers were taken. If the operation is serial, four machines is four times the wait. |