The scenario model, the lifecycle, and what the lab actually costs #6
@@ -37,6 +37,10 @@ Measured, and the finding is a blocker rather than a data point.
|
||||
for one small virtual machine at best — and over two minutes when observed a second time. Cost
|
||||
scales with the number of machines and the size of their disks, not with what changed.
|
||||
|
||||
**With copy-on-write it is 0.13 seconds and costs the delta.** Verified, not assumed. The
|
||||
projected four-machine reset cycle falls from roughly ninety seconds to roughly fifteen, of
|
||||
which almost all is a boot that cannot be avoided.
|
||||
|
||||
The cause is not virtual machines and not incus. It is that the host offers incus exactly one
|
||||
storage driver, `dir`, which has no copy-on-write and therefore no cheap snapshot. The kernel
|
||||
supports btrfs; the userspace tool that would let incus use it is simply not installed.
|
||||
@@ -49,7 +53,8 @@ something the mesh already does.
|
||||
|
||||
| Question | Why it matters |
|
||||
|---|---|
|
||||
| How much does a copy-on-write pool actually improve it? Expected to be near-instant snapshots and delta-sized storage, but **expected is not measured**. | The whole inner-loop argument rests on the answer. |
|
||||
| ~~How much does a copy-on-write pool actually improve it?~~ **Measured: snapshot 9.9 s → 0.13 s, restore 10.4 s → 0.80 s, three snapshots sharing 1.36 GB rather than costing 4.8 GB.** The projected four-machine cycle falls from ~90 s to ~15 s. | Answered. The inner-loop argument holds *with* copy-on-write and did not without it. |
|
||||
| How does the lab install its own prerequisites on a clean machine? | The lab needs a virtualisation daemon, copy-on-write tooling and a pool before it can do anything — and it cannot depend on the mesh for them, since it is where the mesh is built. |
|
||||
| Why was the second snapshot more than twelve times slower than the first? | If snapshot cost is unpredictable rather than merely high, that is worse — a loop with a variable multi-minute step is one nobody trusts. |
|
||||
| Does a scenario snapshot need the machines stopped? | Stateless snapshots of a running virtual machine capture the disk but not memory. Whether a mesh restored that way is coherent is not established. |
|
||||
| What is the cost at scenario scale — four machines rather than one? | Only single-machine numbers were taken. If the operation is serial, four machines is four times the wait. |
|
||||
|
||||
@@ -87,12 +87,55 @@ working capability — which is precisely what was just done for the virtualisat
|
||||
itself, and what [`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)
|
||||
is about.
|
||||
|
||||
## Not measured, and it matters
|
||||
## The fix, measured
|
||||
|
||||
The copy-on-write comparison **was not run**, because running it would mean installing a package
|
||||
by hand, which the mesh's rules forbid and which would have made the measurement unreproducible
|
||||
anyway. Copy-on-write snapshots are expected to be near-instant with storage proportional to
|
||||
what changed. That expectation is well founded and is still an expectation.
|
||||
The comparison was subsequently run. One package — `btrfs-progs`, no dependencies — installed by
|
||||
hand, the daemon restarted so it re-detected drivers, a copy-on-write pool created on a loop
|
||||
file, and the identical image launched onto it.
|
||||
|
||||
Until it is measured, the correct statement is: *the current configuration is too slow, and the
|
||||
likely fix is known but unverified.*
|
||||
| Operation | `dir` | copy-on-write | |
|
||||
|---|---|---|---|
|
||||
| snapshot | 9.9 s, then **> 120 s** | **0.13 s** | ~76× faster, and *consistent* |
|
||||
| snapshot again | — | 0.12 s | |
|
||||
| snapshot a third time | — | 0.13 s | |
|
||||
| restore call | 10.4 s | **0.80 s** | ~13× faster |
|
||||
| restore, to usable | 20.1 s | **10.5 s** | the remainder is boot, which is irreducible |
|
||||
| three snapshots, storage | ~4.8 GB | **1.36 GB total, shared** | cost is the delta, not the disk |
|
||||
|
||||
**The fix is real, and larger than expected.** Snapshot goes from ten seconds to a tenth of a
|
||||
second, and — more importantly — from *wildly variable* to *flat*. Three consecutive snapshots
|
||||
took 0.13, 0.12 and 0.13 seconds. On `dir` the second snapshot never finished.
|
||||
|
||||
Storage stops scaling with the machine and starts scaling with what changed: three snapshots of
|
||||
a 1.5 GB instance occupied 1.36 GB in total, because they share.
|
||||
|
||||
### What it projects to
|
||||
|
||||
A four-machine reset-and-rerun cycle, the operation the inner loop repeats most:
|
||||
|
||||
| | `dir` | copy-on-write |
|
||||
|---|---|---|
|
||||
| snapshot the scenario | ~40 s, 6.4 GB | **~0.5 s**, delta-sized |
|
||||
| restore it | ~40 s + boot | **~3 s** + boot |
|
||||
| **cycle** | **~90 s, unbounded at worst** | **~15 s, dominated by boot** |
|
||||
|
||||
At fifteen seconds, dominated by a boot that cannot be avoided, the inner loop is viable and
|
||||
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)'s argument holds.
|
||||
At ninety it did not.
|
||||
|
||||
### One honest counter-observation
|
||||
|
||||
Launching onto the fresh copy-on-write pool was **slower** — 20.2 s to usable against 14.3 s —
|
||||
because the image had to be unpacked into a pool that had never seen it. That cost is paid once
|
||||
per pool, not per scenario, and it is dwarfed by what snapshotting saves. But it is a real
|
||||
number and it went the other way.
|
||||
|
||||
### State this left behind
|
||||
|
||||
Recorded because hand-made state is exactly what the mesh's rules exist to prevent, and it must
|
||||
be declared properly rather than left as an artefact of a measurement:
|
||||
|
||||
- `btrfs-progs` installed by hand. Its installation regenerated the boot initramfs, a side
|
||||
effect worth knowing about.
|
||||
- The daemon restarted once, to re-detect drivers.
|
||||
- The test pool and instance were **removed**; the pool the lab actually needs does not exist.
|
||||
|
||||
@@ -0,0 +1,116 @@
|
||||
---
|
||||
layer: to-be
|
||||
status: designed
|
||||
code: [mesh-lab]
|
||||
updated: 2026-08-24
|
||||
decisions:
|
||||
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
|
||||
- 02-DECISIONS/0008-a-failed-step-fails-the-job.md
|
||||
---
|
||||
|
||||
# Installing the lab on a clean machine
|
||||
|
||||
The lab has prerequisites — a virtualisation daemon, copy-on-write storage, a pool, an identity
|
||||
permitted to talk to it — and it cannot get them from the mesh, because it is where the mesh is
|
||||
built ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)).
|
||||
|
||||
So the lab needs an install path of its own. This describes it, and the shape it has to take is
|
||||
determined by two failures observed while measuring
|
||||
([research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md)).
|
||||
|
||||
## The two failures that shape this
|
||||
|
||||
**One: installed is not available.** The virtualisation package was present and explicitly
|
||||
installed. Both its units were disabled, the operator was in no group, and the client reported
|
||||
the server unreachable. Nothing had failed — the declaration was satisfied exactly as written
|
||||
([`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)).
|
||||
|
||||
**Two, and worse: available is not adequate.** With the storage tooling absent, the daemon
|
||||
offered one driver, everything worked, and snapshots took **seventy-six times longer** than they
|
||||
needed to. Nothing failed. Nothing warned. A lab in that state runs correctly and is simply too
|
||||
slow to use — and the inner loop it exists to provide quietly does not exist.
|
||||
|
||||
The second is the more dangerous shape and it is a variant the mesh has not catalogued before.
|
||||
Its usual failure is *reported success and did nothing*. This is **reported success and did it
|
||||
seventy-six times slower**, which no error surface catches because nothing is wrong.
|
||||
|
||||
## What follows: the lab verifies capability, never installation
|
||||
|
||||
The install path may differ. **The verification does not.**
|
||||
|
||||
Before the lab will raise anything, it asserts the outcomes it needs — not that packages are
|
||||
present, but that the machine can actually do the work:
|
||||
|
||||
| Assertion | Failing means |
|
||||
|---|---|
|
||||
| the daemon answers **as the invoking user**, not as root | a group membership that was granted but never took effect |
|
||||
| a copy-on-write storage driver is offered | the userspace tooling is missing; snapshots will be full copies |
|
||||
| **the pool the lab will use is on that driver** | a pool exists, and is the slow kind — the failure that has no symptom |
|
||||
| hardware virtualisation is present | machines will be emulated and unusably slow |
|
||||
| an image can be fetched or is cached | the first raise will fail late instead of early |
|
||||
|
||||
Each check states **why it matters**, in the terms of what it costs. *"The pool uses the `dir`
|
||||
driver"* means nothing to someone who does not already know it means seventy-six times slower
|
||||
and unbounded at worst.
|
||||
|
||||
**The lab refuses to run degraded.** It does not warn and continue: a warning about a slow inner
|
||||
loop is read once and ignored forever, and the loop stays slow. This is
|
||||
[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) applied where the failure is
|
||||
performance rather than an error.
|
||||
|
||||
## Two ways the prerequisites arrive
|
||||
|
||||
**On a machine the mesh manages** — a module declares them, and a hook turns them into
|
||||
capabilities: the units enabled, the group granted, the pool created on the right driver. This
|
||||
already works; it is what was done for the virtualisation daemon itself.
|
||||
|
||||
**On a machine the mesh does not manage** — the lab's own bootstrap does it. One command, on a
|
||||
clean machine, that installs what is missing and configures it.
|
||||
|
||||
The second path is not a convenience. It is **required**, because the lab must work before the
|
||||
mesh does, and a lab that could only be installed by a mesh would be unable to host the
|
||||
development of the mesh that installs it.
|
||||
|
||||
Both paths end at the same verification. Whoever satisfied the prerequisites, the lab checks
|
||||
them itself — because the lesson of the first failure is precisely that *something else said it
|
||||
was done* is not evidence.
|
||||
|
||||
## The lab is the second thing installed by hand
|
||||
|
||||
Worth stating, because it looks like an exception and is not.
|
||||
|
||||
The node host is the one thing installed by hand on a machine
|
||||
([research 006](../../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)): everything else
|
||||
arrives through it. The lab is the same shape on a workstation — installed once, by hand, and
|
||||
then everything about the mesh is developed inside it.
|
||||
|
||||
Two bootstraps, at two levels, for the same reason: **something has to be first, and pretending
|
||||
otherwise produces a circularity that gets papered over with a script nobody exercises.**
|
||||
|
||||
## What a clean install actually needs
|
||||
|
||||
In order, on a machine with nothing:
|
||||
|
||||
1. **A virtualisation daemon**, running, with its socket enabled.
|
||||
2. **Copy-on-write storage tooling** — the kernel side is usually already present; it is the
|
||||
userspace half that is missing and that decides whether the driver is offered at all.
|
||||
3. **A pool on that driver.** A loop-backed file is sufficient and needs no partitioning, which
|
||||
matters: requiring a dedicated filesystem would make the lab uninstallable on a machine
|
||||
already in use.
|
||||
4. **Group membership** for the operator — which does **not** apply to sessions that already
|
||||
existed. Observed directly: a shell whose process tree predated the grant could not reach
|
||||
the daemon while a fresh lookup showed the membership present. The bootstrap has to say so,
|
||||
or the first thing a person meets is a permission error that looks like a broken install.
|
||||
5. **Verification**, as above, before anything is raised.
|
||||
|
||||
## Open
|
||||
|
||||
- **Whether the lab's bootstrap may install packages at all**, given that the mesh's rules
|
||||
forbid installing by hand. The resolution is probably that the lab's bootstrap *is* the
|
||||
sanctioned mechanism on an unmanaged machine, in the way the mesh's own first-node script is —
|
||||
but that is an argument to record, not to assume.
|
||||
- **What "adequate" means numerically.** The checks above are qualitative. A snapshot-time
|
||||
threshold would catch a copy-on-write pool that is slow for some other reason, and would be a
|
||||
real assertion rather than a proxy.
|
||||
- **Whether the lab should own its pool** rather than using an existing one. Owning it makes the
|
||||
driver guaranteed; sharing it avoids duplicating storage on a machine that already has a pool.
|
||||
@@ -13,6 +13,7 @@ document is written and this one's status becomes `implemented`.
|
||||
| [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md), [0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) |
|
||||
| [`02-scenario-declaration.md`](02-scenario-declaration.md) | What a scenario declares — the underlay, and what to place on it | [ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md) |
|
||||
| [`03-scenario-lifecycle.md`](03-scenario-lifecycle.md) | What happens to a scenario — raise, snapshot, restore, move, destroy | [ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md) |
|
||||
| [`04-lab-installation.md`](04-lab-installation.md) | Getting the lab onto a clean machine, and why it verifies capability rather than installation | [ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) |
|
||||
|
||||
## Not yet written
|
||||
|
||||
|
||||
Reference in New Issue
Block a user