Files
hq/03-DESIGN/01-to-be
jschoubben 98bcd5cc49 Measure the lab's inner loop — it is too slow, for a fixable reason
The lifecycle design closed on a question that was measurable rather than
arguable, so it was measured. One virtual machine on a workstation with
hardware virtualisation and NVMe.

Raising: the launch call returns in 3.4s, the machine is actually usable
after 14.3s. The gap is a design constraint — raise must wait for the
second number, because reporting the first would be transport reported as
effect, which is the mesh's own recurring failure.

Snapshot: 9.9s and 1.6 GB for a 1.5 GB instance. A dir snapshot is a full
copy; nothing is shared. Restore: 10.4s, usable again after 20.1s.

The second snapshot exceeded two minutes and never completed. That is the
more troubling number: snapshot cost here is not merely high, it is
unpredictable, and a loop with a variable multi-minute step is one nobody
trusts.

Projected to a four-machine scenario, a reset-and-rerun cycle is about a
minute and a half at best and unbounded at worst, before any of the mesh's
own work begins. That is too slow for an inner loop, and ADR 0029's whole
argument — that making the bootstrap path the inner loop turns the
least-exercised code into the most-exercised — holds only while resetting
is cheap.

The cause is not virtual machines. Hardware virtualisation is present and
machines boot in fourteen seconds. It is that the daemon offers exactly one
storage driver, dir, which has no copy-on-write and therefore no cheap
snapshot. The btrfs kernel module is available; btrfs-progs is simply not
installed, which is the entire reason the driver is absent.

The copy-on-write comparison was deliberately NOT run, because running it
would mean installing a package by hand — which the rules forbid and which
would have made the measurement unreproducible. So the honest statement is
that the current configuration is too slow and the likely fix is known but
unverified, rather than that btrfs fixes it.
2026-08-24 00:07:22 +02:00
..

03-DESIGN / 01-to-be

The mesh being built toward. Every statement here traces to a record in 02-DECISIONS/; nothing arrives by drafting.

A document here describes an intention. What currently runs is in 00-as-is/, and the two are never merged — when something ships, the as-is document is written and this one's status becomes implemented.

Document Covers Rests on
00-work-breakdown.md How the decomposition gets built, in what order, and where a human must look ADR 0015
01-end-to-end-testing.md The lab: a real mesh a change can be run against before it reaches nodes ADR 0016, 0029
02-scenario-declaration.md What a scenario declares — the underlay, and what to place on it ADR 0031
03-scenario-lifecycle.md What happens to a scenario — raise, snapshot, restore, move, destroy ADR 0032

Not yet written

  • The eight bounded contexts. ADR 0015 decides the decomposition; the per-context specifications do not exist yet. The work breakdown says in what order they are needed.
  • Domain grouping outside the core. ADR 0017 settles the principle and explicitly does not settle the domain list. That is a research effort, not a design document, until it concludes.