Jochen: a normal application has 3-5 ADRs, maybe 10 for a large one, and we are at 65. Fair, and the cause is mine -- I recorded every FINDING as a decision rather than every fork in the road. Two merges, both cases where one decision had been split across many records because it was taken over several days rather than at once. 0019 absorbs ten records about how this repository works: what it is and that it is public, the folder flow, the two design layers, the issue front door, status in frontmatter, playbooks, the naming rule, the product name. Those were never ten decisions -- they were one, seen from ten angles as the repository took shape. 0016 absorbs the five about the lab: a node is a virtual machine, a router is scenery, a scenario declares the underlay, a scenario is a closed address space, and the two scenario classes. Same pattern -- one design, split by the order it was worked out in. The consolidated 0019 also raises the bar for what earns a record, since that is what produced 65: a record is warranted when there is a genuine fork -- a direction reversed, an alternative that will be proposed again, something contested. A finding is not a decision, and a bug is certainly not. Everything else belongs in the design document where the reasoning is actually read. The checker earned its place here. Deleting nine records left 13 dangling links across the repository and it named every one, including in AGENTS.md. Nothing was found by reading. Remaining clusters worth the same treatment: the host (8 records), delivery (5), modules (6), connectivity (4), substrate and control plane (4). That would be 52 down to roughly 30.
3.4 KiB
status, initiated, touches, became
| status | initiated | touches | became | |||
|---|---|---|---|---|---|---|
| active | 2026-08-24 |
|
010 — What the lab's inner loop actually costs
What is being investigated
The lifecycle design closes on an open question that is measurable rather than arguable:
What a snapshot costs. Whole-scenario snapshots of several machines are the operation the inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop is slow, and everything above is theory.
Measured on a workstation, 2026-08-24. Numbers in measurements.md.
Why it matters
ADR 0016 makes the bootstrap scenario the inner development loop for tiers 0 and 1 — the argument being that raising a node from nothing stops being the least-exercised path and becomes the most-exercised one. That argument is only true if raising and resetting are cheap. A loop that costs minutes is a loop people avoid, and the least-exercised path stays least-exercised.
Status
Measured, and the finding is a blocker rather than a data point.
A snapshot on the current host is a full copy of the machine's disk. 1.6 GB and ten seconds for one small virtual machine at best — and over two minutes when observed a second time. Cost scales with the number of machines and the size of their disks, not with what changed.
With copy-on-write it is 0.13 seconds and costs the delta. Verified, not assumed. The projected four-machine reset cycle falls from roughly ninety seconds to roughly fifteen, of which almost all is a boot that cannot be avoided.
The cause is not virtual machines and not incus. It is that the host offers incus exactly one
storage driver, dir, which has no copy-on-write and therefore no cheap snapshot. The kernel
supports btrfs; the userspace tool that would let incus use it is simply not installed.
So the question "is the lab's inner loop fast enough" currently answers itself the wrong way, for a reason that is one declared package away from being fixed — and declaring packages is something the mesh already does.
Open questions
| Question | Why it matters |
|---|---|
| Answered. The inner-loop argument holds with copy-on-write and did not without it. | |
| How does the lab install its own prerequisites on a clean machine? | The lab needs a virtualisation daemon, copy-on-write tooling and a pool before it can do anything — and it cannot depend on the mesh for them, since it is where the mesh is built. |
| Why was the second snapshot more than twelve times slower than the first? | If snapshot cost is unpredictable rather than merely high, that is worse — a loop with a variable multi-minute step is one nobody trusts. |
| Does a scenario snapshot need the machines stopped? | Stateless snapshots of a running virtual machine capture the disk but not memory. Whether a mesh restored that way is coherent is not established. |
| What is the cost at scenario scale — four machines rather than one? | Only single-machine numbers were taken. If the operation is serial, four machines is four times the wait. |