diff --git a/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md b/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md new file mode 100644 index 0000000..db40f90 --- /dev/null +++ b/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md @@ -0,0 +1,60 @@ +--- +status: active +initiated: 2026-08-24 +touches: + - 03-DESIGN/01-to-be/03-scenario-lifecycle.md + - 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md + - 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md +became: [] +--- + +# 010 — What the lab's inner loop actually costs + +## What is being investigated + +[The lifecycle design](../../03-DESIGN/01-to-be/03-scenario-lifecycle.md) closes on an open +question that is measurable rather than arguable: + +> **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the +> inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop +> is slow, and everything above is theory. + +Measured on a workstation, 2026-08-24. Numbers in [`measurements.md`](measurements.md). + +## Why it matters + +[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) makes the +bootstrap scenario the inner development loop for tiers 0 and 1 — the argument being that +raising a node from nothing stops being the least-exercised path and becomes the most-exercised +one. **That argument is only true if raising and resetting are cheap.** A loop that costs +minutes is a loop people avoid, and the least-exercised path stays least-exercised. + +## Status + +Measured, and the finding is a blocker rather than a data point. + +**A snapshot on the current host is a full copy of the machine's disk.** 1.6 GB and ten seconds +for one small virtual machine at best — and over two minutes when observed a second time. Cost +scales with the number of machines and the size of their disks, not with what changed. + +**With copy-on-write it is 0.13 seconds and costs the delta.** Verified, not assumed. The +projected four-machine reset cycle falls from roughly ninety seconds to roughly fifteen, of +which almost all is a boot that cannot be avoided. + +The cause is not virtual machines and not incus. It is that the host offers incus exactly one +storage driver, `dir`, which has no copy-on-write and therefore no cheap snapshot. The kernel +supports btrfs; the userspace tool that would let incus use it is simply not installed. + +So the question *"is the lab's inner loop fast enough"* currently answers itself the wrong way, +for a reason that is one declared package away from being fixed — and declaring packages is +something the mesh already does. + +## Open questions + +| Question | Why it matters | +|---|---| +| ~~How much does a copy-on-write pool actually improve it?~~ **Measured: snapshot 9.9 s → 0.13 s, restore 10.4 s → 0.80 s, three snapshots sharing 1.36 GB rather than costing 4.8 GB.** The projected four-machine cycle falls from ~90 s to ~15 s. | Answered. The inner-loop argument holds *with* copy-on-write and did not without it. | +| How does the lab install its own prerequisites on a clean machine? | The lab needs a virtualisation daemon, copy-on-write tooling and a pool before it can do anything — and it cannot depend on the mesh for them, since it is where the mesh is built. | +| Why was the second snapshot more than twelve times slower than the first? | If snapshot cost is unpredictable rather than merely high, that is worse — a loop with a variable multi-minute step is one nobody trusts. | +| Does a scenario snapshot need the machines stopped? | Stateless snapshots of a running virtual machine capture the disk but not memory. Whether a mesh restored that way is coherent is not established. | +| What is the cost at scenario scale — four machines rather than one? | Only single-machine numbers were taken. If the operation is serial, four machines is four times the wait. | diff --git a/01-RESEARCH/010-lab-inner-loop-cost/measurements.md b/01-RESEARCH/010-lab-inner-loop-cost/measurements.md new file mode 100644 index 0000000..a690054 --- /dev/null +++ b/01-RESEARCH/010-lab-inner-loop-cost/measurements.md @@ -0,0 +1,141 @@ +--- +effort: 010-lab-inner-loop-cost +updated: 2026-08-24 +--- + +# Measurements + +Taken 2026-08-24 on a workstation with hardware virtualisation available, an NVMe-backed ext4 +root, and 300 GB free. One virtual machine, 1 GiB memory, 2 CPUs, from a cached distribution +image. + +## The environment, before anything ran + +| Fact | Value | Consequence | +|---|---|---| +| Hardware virtualisation | present | virtual machines run at native speed; the choice in [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md) is not paying an emulation penalty | +| Storage drivers the daemon offers | **`dir` only** | no copy-on-write, therefore no cheap snapshot | +| Host filesystems | ext4 throughout | nothing copy-on-write to put a pool on | +| btrfs kernel module | **available** | the kernel can do it | +| `btrfs-progs` | **not installed** | which is the entire reason the driver is absent | + +The last two rows are the finding. The daemon advertises only `dir` because the userspace tool +for anything better is missing — not because the host cannot do better. + +## Raising a machine + +| Step | Time | +|---|---| +| launch call returns | **3.4 s** | +| machine actually usable — a command executes on it | **14.3 s** | + +The gap matters for the lifecycle design: `raise` returning is not the same as the scenario +being ready, so the verb has to wait for the second number, not report the first. Reporting +the first would be the mesh's own recurring failure — transport reported as effect. + +## Snapshot and restore + +| Operation | Time | Disk | +|---|---|---| +| snapshot, first | **9.9 s** | **+1.6 GB** | +| snapshot, second | **> 120 s — did not complete** | — | +| restore call returns | **10.4 s** | — | +| machine usable again | **20.1 s** total | — | + +Instance on disk before snapshotting: 1.5 GB. Snapshot directory afterwards: 1.6 GB. **A `dir` +snapshot is a full copy** — the storage cost equals the instance, and nothing is shared. + +Implied copy throughput on the first snapshot is roughly 160 MB/s, which is far below what the +underlying NVMe can do and is consistent with a real, durable copy rather than a metadata +operation. + +**The second snapshot is the more troubling number.** It exceeded two minutes and was still +running when the observation was cut off; only the first snapshot exists. Whatever the cause — +page cache exhausted by the preceding restore, writeback contention — the practical +consequence is that snapshot cost here is **not merely high, it is unpredictable**. + +## What this projects to + +A four-machine scenario, taking the optimistic single-machine numbers and assuming the +operations are serial: + +| | one machine | four machines | +|---|---|---| +| raise, to usable | 14 s | ~57 s | +| snapshot | 10 s, 1.6 GB | ~40 s, 6.4 GB | +| restore, to usable | 20 s | ~80 s | + +A reset-and-rerun cycle is therefore **around a minute and a half at best**, and unbounded at +worst, before any of the mesh's own work begins. + +## The judgement + +**This is too slow for an inner loop**, and the reason is not the design. + +[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) argues that +making the bootstrap path the inner development loop turns the least-exercised code in the +system into the most-exercised. That argument holds only while resetting is cheap. At a minute +and a half a cycle, with occasional multi-minute stalls, the loop is one a person works around +— and the path stays under-exercised for exactly the reason it always was. + +Nothing about virtual machines causes this. Hardware virtualisation is present and the machines +boot in fourteen seconds. **The cost is entirely the storage driver**, and the driver is absent +because one userspace package is not installed on the host. + +The mesh already has the mechanism for that: a module declares a package, and a hook makes it a +working capability — which is precisely what was just done for the virtualisation daemon +itself, and what [`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md) +is about. + +## The fix, measured + +The comparison was subsequently run. One package — `btrfs-progs`, no dependencies — installed by +hand, the daemon restarted so it re-detected drivers, a copy-on-write pool created on a loop +file, and the identical image launched onto it. + +| Operation | `dir` | copy-on-write | | +|---|---|---|---| +| snapshot | 9.9 s, then **> 120 s** | **0.13 s** | ~76× faster, and *consistent* | +| snapshot again | — | 0.12 s | | +| snapshot a third time | — | 0.13 s | | +| restore call | 10.4 s | **0.80 s** | ~13× faster | +| restore, to usable | 20.1 s | **10.5 s** | the remainder is boot, which is irreducible | +| three snapshots, storage | ~4.8 GB | **1.36 GB total, shared** | cost is the delta, not the disk | + +**The fix is real, and larger than expected.** Snapshot goes from ten seconds to a tenth of a +second, and — more importantly — from *wildly variable* to *flat*. Three consecutive snapshots +took 0.13, 0.12 and 0.13 seconds. On `dir` the second snapshot never finished. + +Storage stops scaling with the machine and starts scaling with what changed: three snapshots of +a 1.5 GB instance occupied 1.36 GB in total, because they share. + +### What it projects to + +A four-machine reset-and-rerun cycle, the operation the inner loop repeats most: + +| | `dir` | copy-on-write | +|---|---|---| +| snapshot the scenario | ~40 s, 6.4 GB | **~0.5 s**, delta-sized | +| restore it | ~40 s + boot | **~3 s** + boot | +| **cycle** | **~90 s, unbounded at worst** | **~15 s, dominated by boot** | + +At fifteen seconds, dominated by a boot that cannot be avoided, the inner loop is viable and +[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)'s argument holds. +At ninety it did not. + +### One honest counter-observation + +Launching onto the fresh copy-on-write pool was **slower** — 20.2 s to usable against 14.3 s — +because the image had to be unpacked into a pool that had never seen it. That cost is paid once +per pool, not per scenario, and it is dwarfed by what snapshotting saves. But it is a real +number and it went the other way. + +### State this left behind + +Recorded because hand-made state is exactly what the mesh's rules exist to prevent, and it must +be declared properly rather than left as an artefact of a measurement: + +- `btrfs-progs` installed by hand. Its installation regenerated the boot initramfs, a side + effect worth knowing about. +- The daemon restarted once, to re-detect drivers. +- The test pool and instance were **removed**; the pool the lab actually needs does not exist. diff --git a/02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md b/02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md new file mode 100644 index 0000000..09dac33 --- /dev/null +++ b/02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md @@ -0,0 +1,71 @@ +--- +status: accepted +date: 2026-08-23 +deciders: jochen +reconstructed: false +--- + +# 32. A scenario is an isolated address space, and the lab never reaches into it over IP + +## Context + +A scenario declares literal addresses — +[the declaration](../03-DESIGN/01-to-be/02-scenario-declaration.md) is full of them, and it has +to be, because reproducing *published but behind NAT* means saying which address the world sees. + +That raises a question the declaration left open: **two scenarios at once.** Several agents +working means several scenarios, and the lab design already calls that a requirement. But two +scenarios built from the same declaration want the same addresses, and there are only three +documentation ranges in existence. + +## Considered options + +1. **Allocate addresses from a pool at raise time**, rewriting the declaration's literals. + Rejected. It makes the addresses in a declaration a fiction, so a scenario reproducing a + specific topology no longer reproduces it; it breaks the RFC-range validation, since + allocated addresses would have to come from somewhere real; and the numbers a person reads + in the file stop being the numbers they will see in a capture. +2. **One scenario at a time.** Rejected — it is the requirement, not an inconvenience. A gate + an agent has to queue for is a gate that gets bypassed. +3. **Give each scenario its own network stack, so the addresses do not collide.** Chosen. + +## Decision + +**A scenario is a closed address space.** Every segment materialises as its own isolated link, +belonging to one scenario instance. Two scenarios raised from the same declaration hold the same +addresses and never meet, because nothing joins their links. + +The declaration therefore keeps its literal addresses, and they mean exactly what they say. + +**The consequence that constrains everything else: the lab never reaches into a scenario over +IP.** It talks to a machine through the virtualisation layer's own channel — the same way one +executes a command in a container without the container being routable. + +That is not a preference. If the lab reached machines by address, the workstation running it +would need a route into each scenario, and two scenarios carrying the same prefix would give it +two routes to the same destination. Concurrency would be impossible, and it would fail in the +worst available way: not with an error, but by one scenario's traffic arriving in another. + +## Consequences + +- Scenarios are concurrent by construction, with no allocation, no bookkeeping and no limit + beyond the machine's capacity. +- The three documentation ranges stop being a scarce resource. Every scenario may use all of + them, because no two scenarios share a link. +- **The lab cannot use IP to check anything**, which is more of a constraint than it first + appears: *"can this machine reach that one"* has to be asked **from inside the scenario**, by + executing on a machine, rather than probed from outside. That is the honest way to ask it + anyway — reachability from the workstation is not the question. +- A scenario is a unit that can be paused, snapshotted and destroyed whole, because nothing + outside holds a reference into it. +- The lab needs a scenario **instance** identity distinct from the scenario name in the + declaration: the declaration is a kind, and several instances of one kind may exist. +- **The workstation is not on the scenario's network, so it is not a node in it.** Anything a + developer wants to reach — a web interface, a database — needs an explicit, deliberate + forward out of the scenario, which is a feature rather than a gap: nothing leaks by default. + +## References + +- [ADR 0031](0031-the-lab-provides-the-underlay.md) — the declaration whose literal addresses + this preserves. +- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the lifecycle jobs this shapes. diff --git a/03-DESIGN/01-to-be/02-scenario-declaration.md b/03-DESIGN/01-to-be/02-scenario-declaration.md index a159a3a..370be64 100644 --- a/03-DESIGN/01-to-be/02-scenario-declaration.md +++ b/03-DESIGN/01-to-be/02-scenario-declaration.md @@ -17,31 +17,191 @@ everything in the lab hangs off, so it is worth getting small. It states what a hosting provider and a home router would provide, and nothing the mesh is responsible for ([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)). +## Public networks are unrelated, and routed rather than bridged + +The internet is not a network. It is a very large number of unrelated networks that route to +each other, and a machine in one is **many hops** from a machine in another with no shared +broadcast domain between them. + +So a scenario does not have *an* internet segment. It has **one public segment per public +network**, each with its own unrelated prefix, and the lab wires them together **through a +router, never onto a shared bridge**. + +That distinction is load-bearing, and putting several public addresses in one prefix would +quietly make four things true that are false in reality: + +| If public addresses share a segment | Reality | +|---|---| +| machines resolve each other by ARP and talk directly | they are routed, many hops apart | +| TTL never decrements | every hop decrements it | +| broadcast and multicast reach across | neither crosses a router | +| any two are adjacent | adjacency is the exception, not the rule | + +The third is not hypothetical here. The mesh has already been bitten by multicast name +resolution — [`00-as-is/01-mesh-and-transport.md`](../00-as-is/01-mesh-and-transport.md) +records that mesh names are deliberately not multicast names, after a delay and a +one-node-only failure mode. A lab where "the internet" is one broadcast domain would let a node +discover a peer by multicast that it could never discover in production, and report success. + +**The rule: `kind: public` segments are routed to one another and never bridged.** It is a +property of how the lab wires them, not a field anyone sets, because there is no correct +scenario in which two public networks are adjacent. + +### Which addresses to use + +RFC 5737 reserves three ranges and RFC 3849 reserves one v6 prefix. Each public network takes +its own, and they are chosen to look nothing like each other — because in reality they would +not: + +| Public network | v4 | v6 | +|---|---|---| +| a hosting provider | `192.0.2.0/24` | `2001:db8:a::/48` | +| a household ISP | `198.51.100.0/24` | `2001:db8:b::/48` | +| a mobile or foreign network | `203.0.113.0/24` | `2001:db8:c::/48` | + +Three is enough for the topologies that matter, and a fourth public network can subnet one of +them — an ISP handing out `198.51.100.0/25` and `198.51.100.128/25` to two customers is exactly +what really happens. + +Private segments still use RFC 1918 and stay byte-identical to production. + +## What NAT does, and why the design turns on it + +A household or office has **one** address the outside world can see, and **many** machines +behind it. Network address translation is what reconciles those. + +When a machine inside dials out, the gateway rewrites the packet's source from the private +address to the public one, **remembers the mapping**, and rewrites the replies on the way back. +Four consequences follow, and every one of them shapes this design: + +1. **Outbound works; inbound does not.** A mapping exists only because something inside started + a conversation. Nothing outside can start one — there is no mapping to look up, and no way + to know which internal machine was meant. +2. **A forwarded port is a permanent mapping made by hand**, in the inbound direction: + *anything arriving at the public address on 443 goes to this machine.* That is the only way + a machine behind NAT becomes reachable, and it requires control of the gateway. +3. **Mappings expire.** A gateway forgets one that goes unused. This is why anything holding a + connection through NAT sends keepalives, and why a mesh that does not is fine until it is + idle. +4. **From outside, every machine behind the gateway looks like one address.** Identity and + address stop corresponding. + +This is why the mesh dials outward and never inward +([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md)), why a hub exists at +all, and why a node's endpoint is something a peer **learns** from arriving packets rather than +something anyone configures. + +**Carrier-grade NAT** is the same mechanism applied by an ISP: your own gateway gets a private +address too, and the public one is shared with strangers. Nothing can be forwarded, because the +rule would have to live on equipment you do not own. Common on mobile connections and +increasingly on fixed ones. + +## Three positions a machine can be in + +The underlay's whole job is to reproduce **where a machine sits relative to the internet**, +because that is what the mesh has to cope with and what only production currently exercises. +There are three positions, and they are genuinely different: + +| Position | Reachable from outside | Apparent address | Example | +|---|---|---|---| +| **Attached** | yes, at its own address | its own | a hosted server | +| **Behind a forwardable gateway** | only through a forwarded port, at the *gateway's* address | the gateway's | a machine at home | +| **Behind an unforwardable gateway** | **no** | someone else's, and it changes | a laptop on a café network; anything behind carrier-grade NAT | + +The axis is **forwardability, not ownership** — which is worth stating because the obvious +framing gets it wrong. Carrier-grade NAT is *your* connection and is still unforwardable, so it +belongs in the third row alongside the café. What the mesh has to cope with is whether an +inbound mapping can be made, not who owns the equipment. + +The third position is the hard one. A machine there can dial out and nothing more: it cannot be +published, its apparent address belongs to a router it does not control, and that address +changes when it moves. Every assumption a mesh makes about reachability breaks there first. + +A declaration has to be able to say all three, and to move a machine between them. + +## Reachability is per address family, not per machine + +Adding IPv6 is not a field. It changes the position model, and the reason is worth stating +before the syntax. + +**IPv6 usually has no NAT.** A machine behind a household gateway can hold a *globally routable* +v6 address while its v4 address is private and unforwardable. The same machine, at the same +moment, is in **two different positions at once**: + +| | IPv4 | IPv6 | +|---|---|---| +| a typical machine at home | behind an unforwardable-or-forwardable gateway | **attached**, directly reachable | +| a machine on mobile data | behind carrier-grade NAT | often attached, sometimes absent entirely | +| a machine on an older network | attached or behind NAT | **no address at all** | + +So the three positions apply **per family**, and a machine's reachability is a property of +*(machine, family)* rather than of the machine. A mesh that treats reachability as one fact per +node will reach a peer over one family, fail over the other, and report whichever it tried. + +That has a direct consequence for what the lab is for: *"can these two nodes reach each other"* +stops being a yes/no question. It is asked once per family, and the interesting answers are the +asymmetric ones. + +The v6 documentation prefix is `2001:db8::/32` (RFC 3849) — the exact counterpart of the RFC +5737 rule, and load-bearing for the same reason. + ## The shape ```yaml -scenario: published-behind-nat +scenario: roaming-and-published segments: - wan: - cidr: 203.0.113.0/24 # RFC 5737 — never routes on the real internet - lan: - cidr: 192.168.1.0/24 - behind: wan # NAT; the lab materialises a router + hosting: # one public network + kind: public + cidr: [192.0.2.0/24, 2001:db8:a::/48] + + isp-home: # another, unrelated — routed to it, not bridged + kind: public + cidr: [198.51.100.0/24, 2001:db8:b::/48] + + home: + kind: private + cidr: [192.168.1.0/24, 2001:db8:b:1::/64] + mtu: 1500 + gateway: + to: isp-home + address: [198.51.100.7, 2001:db8:b::7] # what the world sees this network as + nat: [v4] # v4 is translated; v6 is routed, not translated + forwardable: true + mapping_ttl: 120s # an unused inbound mapping is forgotten after this + + cafe: # a network we do not control + kind: private + cidr: [10.50.0.0/16] # v4 only — no v6 offered here at all + mtu: 1400 # a tunnelled path, smaller than standard + gateway: + to: hosting # it reaches the world via a different public network + address: [192.0.2.200] + nat: [v4] + forwardable: false # café wifi, or carrier-grade NAT + mapping_ttl: 30s # aggressive, as carrier NAT tends to be machines: anchor: - segment: wan - address: 203.0.113.10 + at: { segment: hosting, address: [192.0.2.10, 2001:db8:a::10] } + home-server: - segment: lan - address: 192.168.1.135 - forwarded: [443] # reachable from wan through the router + at: { segment: home, address: [192.168.1.135, 2001:db8:b:1::135] } + published: + - { port: 443, on: home } # v4 only: 198.51.100.7:443 → 192.168.1.135:443 + inbound: allow # v6 is routable here, so this decides whether it is reachable + workstation: - segment: lan - address: 192.168.1.250 + at: { segment: home, address: [192.168.1.250, 2001:db8:b:1::250] } + inbound: deny # a host firewall: dials out, accepts nothing + laptop: - segment: detached # reachable by nothing until attached + at: { segment: home, address: [192.168.1.98, 2001:db8:b:1::98] } + + border: # a machine on two segments at once + at: + - { segment: home, address: [192.168.1.2] } + - { segment: isp-home, address: [198.51.100.60] } place: all: [host] @@ -50,41 +210,113 @@ place: snapshot: raised ``` -That is a complete bootstrap scenario. Nothing in it mentions the overlay, a hub, peering, -names or certificates — all of which are outcomes to be observed. +## What each part means, precisely -## The four parts +**`segments`** — a broadcast domain with an address range, and a `kind:`. A segment is a single +broadcast domain, which is precisely why several public networks cannot be one segment. -**`segments`** — the networks that exist. `behind:` declares NAT, and is the only place a -router comes from: the lab materialises one without being asked, because NAT has to run -somewhere. This is the one implicit machine in an otherwise explicit declaration. +`kind: public` marks a segment that stands in for a public network — and there is normally more +than one, unrelated to each other. `kind: private` is everything else. This is stated rather than inferred, and the earlier version inferred it — *a segment with +no gateway is the internet* — which made an **isolated network inexpressible**: a LAN with no +route out is a private segment with no gateway, and would have been read as public and forced +to use documentation addresses. A mesh spanning a site with no internet access is a real +topology, and the model has to be able to say it. -**`machines`** — what sits where. A machine has a segment and an address, and that is nearly -all. `forwarded:` opens a port through the router, which is what makes *published but behind -NAT* reproducible — the case that exists only in production today. `segment: detached` is a -machine on no network, which is how a roaming node is expressed at rest. +**`gateway:`** — how a segment reaches its parent, and this is where the previous version was +too thin. It carries three facts, and all three are load-bearing: -**`place`** — what goes inside. `all:` applies to every machine; a machine name overrides for -that machine. This is the only part that differs between the two scenario classes. +- `to:` — the parent segment. +- `address:` — **the address the outside world sees this network as.** For a household + connection this is the public address the ISP hands out. It is not decoration: it is what a + peer records as the endpoint when a machine here dials out, and what a public name for a + published machine here resolves to. +- `nat:` — **which families are translated**, as a list. `[v4]` is the ordinary modern case: + v4 translated, v6 routed. `[v4, v6]` describes a gateway that translates both, which exists + and is worth being able to reproduce. `[]` is a routed range, where machines keep their own + addresses and the gateway only forwards. +- `forwardable:` — whether an inbound mapping can be created. Independent of `nat:`, and the + field that separates a home gateway from carrier-grade NAT. Publishing through a gateway with + `forwardable: false` is a declaration error, because that is exactly the constraint being + reproduced. +- `mapping_ttl:` — how long an unused inbound mapping survives. This is what makes keepalive + behaviour testable: a mesh that holds a connection through NAT without refreshing it works + perfectly until the far side goes quiet for longer than this. Aggressive values reproduce + carrier NAT; omitting it means mappings never expire, which no real gateway does. -**`snapshot`** — names the state once placement finishes, so a run can return to it without -raising everything again. Snapshots are what make repetition cheap, and cheap repetition is -what makes the bootstrap path the inner development loop rather than a ceremony. +**`segments[].mtu`** — the largest packet the segment carries, defaulting to 1500. Lower values +reproduce tunnelled and PPPoE paths. This matters because an overlay adds its own header: a +tunnel over a 1400-byte path establishes a connection and then silently drops large packets, +which is the shape of fault this whole effort exists to stop shipping. + +**`machines[].inbound`** — `allow` or `deny`, a host firewall. Distinct from NAT and behaves +differently: a machine can be perfectly routable and still refuse everything unsolicited, which +is the normal state of a v6-addressed machine. Without this, v6 addressing would imply +reachability, and it does not. + +The lab materialises a machine to be the gateway. That is the one implicit machine in an +otherwise explicit declaration, and it exists because NAT has to run somewhere. + +**`machines[].at`** — segment and addresses, or a **list** of them for a machine on several +segments at once. Multi-homing is not exotic: it is what a border machine is, and what any node +with both a LAN and a WAN interface is. Each entry carries the addresses that machine holds on +that segment, one per family. + +Position follows from the pair, per family: on a `kind: public` segment a machine is attached; +on a private one it is behind that segment's gateway, unless the gateway does not translate +that family — in which case it is attached on that family and behind a gateway on the other. + +**`machines[].published`** — a destination-NAT rule on a named gateway, stated as an outcome +rather than a port list. `{ port: 443, on: home }` means the `home` gateway forwards its own +`203.0.113.50:443` to this machine's `443`. The resulting public endpoint is derivable, which is +the point: a scenario never writes an endpoint down, and the mesh has to discover it. + +A machine may be published on **any gateway between it and a public network** — which is how *"our +LAN also has a public IP"* is expressed, and why `on:` names the gateway rather than being +implied. It cannot be published at all on a gateway the scenario models as foreign; attempting +it is a declaration error, because that is precisely the constraint being reproduced. + +**`at: detached`** — on no segment. A machine that exists and can reach nothing. + +## Moving a machine is a lifecycle operation + +`at:` states where a machine *starts*. Moving it is something a run does: + +``` +move laptop → { segment: elsewhere, address: 198.51.100.23 } +move laptop → detached +move laptop → { segment: home, address: 192.168.1.98 } +``` + +This is the roaming case made testable, and it is the one that finds the interesting faults. +The same machine, the same identity, three positions in one run: at home where its peers can +reach it directly, on a foreign network where it can only dial out and its apparent address +belongs to a router it does not control, and asleep. + +Whether the overlay survives that, re-forms, and is noticed to have changed endpoint is +**observed**, never arranged +([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)). ## Why the addresses are load-bearing -The routable segment uses RFC 5737 documentation space, and this is not a stylistic choice. +The internet segment uses RFC 5737 documentation space, and this is not a stylistic choice. The mesh decides *public versus private* by matching the address. A private range on the -segment meant to be routable makes the hub test as unreachable, and **the mesh silently never -forms** — no error, no failed step, just a mesh that does not exist. Research 004 calls this -the single most important fact in its analysis. +segment meant to be routable makes a would-be hub test as unreachable, and **the mesh silently +never forms** — no error, no failed step, just a mesh that does not exist. Research 004 calls +this the single most important fact in its analysis. -So the format should make this hard to get wrong rather than merely documented: a segment -without `behind:` is a routable segment, and an address in it that is not documentation space -is a declaration error, refused before anything is raised. That is +Which range goes where is covered above, under *public networks are unrelated*: one range per +public network, chosen to look nothing like each other. + +Private segments use RFC 1918 and can be **byte-identical to production**, because those +addresses mean the same thing everywhere. Only the public side is substituted, and only because +it must be. + +The format should make getting this wrong hard rather than merely documented: a segment without +a `gateway:` is a public segment, and an address in it — including a gateway's `address:` — that +is not documentation space is a declaration error, refused before anything is raised. That is [ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) applied to a configuration -file — the failure it prevents is silent, so the check has to be loud. +file: the failure it prevents is silent, so the check has to be loud. ## The same declaration serves both classes @@ -124,18 +356,284 @@ not first. - **Steps.** A scenario is a desired state. Anything expressed as an ordered list of actions belongs in the lifecycle, not the declaration. +## What the model deliberately does not contain + +A real network is full of equipment: a modem, a router, switches, access points, controllers. +Almost none of it appears here, and the omission is deliberate rather than an oversight. + +**The test is whether a device changes what an IP packet can do.** If two machines can exchange +packets, at the same addresses, with the same reachability and the same MTU, whether or not the +device exists — then the device is invisible to the mesh, and modelling it would add a fixture +with no fault to catch. + +| Equipment | Modelled? | Why | +|---|---|---| +| **Switch** | no | Moves frames within a segment. Two machines on a switch are two machines on a segment. | +| **Access point** | no | Bridges wireless clients onto a segment. A machine on wifi and a machine on cable are the same machine to IP. | +| **Network controller** | no | Configures equipment. Its effects appear as segments and policy; it has no packets of its own. | +| **Cabling, PoE, uplink speed** | no | Change performance and availability, not reachability. | +| **Router / security gateway** | **yes — it *is* the gateway** | Translation, forwarding and inter-segment policy all live here. | +| **ISP modem** | **only in router mode** | In bridge mode it is a media converter and invisible. Doing its own NAT, it is a second gateway — and that is double NAT. | +| **VLANs** | **yes — they are segments** | Machines on separate VLANs cannot reach each other except through the router, which is the definition of a separate segment. | +| **Inter-VLAN firewall rules** | **yes** — see below | A rule stopping one segment reaching another is a reachability fact, and the mesh will hit it. | + +The two entries worth dwelling on are the ones where a common household setup produces +something the mesh has to survive. + +**A modem in router mode gives you double NAT.** Your gateway holds a private address from the +modem, which holds the public one. Forwarding then requires a rule on *both*, and one of them +may not be configurable. This is expressible as nested segments — `to:` chains — but publishing +through two gateways is not, and remains open. + +**Segmented networks are the common case, not the exotic one.** A router with separate networks +for trusted machines, guests and devices is ordinary, and the rules between them are ordinary +too. A mesh node on one segment and a mesh node on another are, as far as reachability goes, on +different networks that happen to share a gateway. + +## Segments may share a gateway + +Several segments can name the same parent and the same external address. That is one router +with several networks behind it, which is what a VLAN-capable gateway is: + +```yaml +segments: + trusted: + kind: private + cidr: [192.168.1.0/24] + gateway: { to: isp-home, address: [198.51.100.7], nat: [v4], forwardable: true } + devices: + kind: private + cidr: [192.168.30.0/24] + gateway: { to: isp-home, address: [198.51.100.7], nat: [v4], forwardable: true } +``` + +Identical gateway declarations mean **one gateway machine**, not two. The lab materialises a +single router serving both, because that is what the topology being reproduced is — and two +routers sharing one address would not work anyway. + +## Policy between segments + +Sharing a gateway does not mean segments can reach each other. What they may do is stated +separately, because it is a fact about a pair rather than a property of either: + +```yaml +policy: + - { from: devices, to: trusted, allow: false } # devices may not initiate to trusted + - { from: trusted, to: devices, allow: true } # the reverse is fine +``` + +Default is `true` between segments behind the same gateway, matching a router with no rules +configured. Asymmetry is the normal case and the reason this cannot be a single flag: the +useful configuration is almost always one-directional. + +This is **not** the same as `machines[].inbound`, and conflating them loses a real distinction: + +| | Enforced by | Blocks | +|---|---|---| +| `policy` | the gateway, between segments | everything crossing, regardless of what the destination thinks | +| `inbound` | the machine itself | unsolicited traffic that already reached it | + +A mesh node behind a `policy` deny cannot be reached even by a peer that knows exactly where it +is — and that is a real topology, not a contrived one. + +## Worked example — a whole mesh of the ordinary kind + +The shape research 004 identified: one machine with a routable address, one publicly named but +behind a household connection, one stationary machine on that network, one that roams. Written +out completely, with every field the model has. + +```yaml +scenario: the-ordinary-shape + +segments: + # ---- three unrelated public networks. Routed to each other, never bridged. ---- + + hosting: # where the always-on machine lives + kind: public + cidr: [192.0.2.0/24, 2001:db8:a::/48] + + isp-home: # the household's uplink + kind: public + cidr: [198.51.100.0/24, 2001:db8:b::/48] + + isp-mobile: # wherever the roaming machine happens to be + kind: public + cidr: [203.0.113.0/24, 2001:db8:c::/48] + + # ---- private networks behind them ---- + + home: # the household network + kind: private + cidr: [192.168.1.0/24, 2001:db8:b:1::/64] + mtu: 1492 # PPPoE on the uplink; 1500 if the line is not PPPoE + gateway: + to: isp-home + address: [198.51.100.7, 2001:db8:b::7] + nat: [v4] # v4 translated, v6 routed — the modern default + forwardable: true # the household router is ours to configure + mapping_ttl: 120s + + devices: # optional: a segmented network on the same router + kind: private + cidr: [192.168.30.0/24] + gateway: + to: isp-home + address: [198.51.100.7, 2001:db8:b::7] # identical → the SAME gateway machine + nat: [v4] + forwardable: true + mapping_ttl: 120s + + cafe: # a network we do not control + kind: private + cidr: [10.50.0.0/16] # RFC 1918 — someone else's private range + mtu: 1400 + gateway: + to: isp-mobile + address: [203.0.113.129] + nat: [v4] + forwardable: false # carrier-grade NAT, or simply not ours + mapping_ttl: 30s + +policy: + - { from: devices, to: home, allow: false } + - { from: home, to: devices, allow: true } + +machines: + anchor: # routable, nothing in front of it + at: { segment: hosting, address: [192.0.2.10, 2001:db8:a::10] } + inbound: allow + + home-server: # publicly named, behind the household connection + at: { segment: home, address: [192.168.1.135, 2001:db8:b:1::135] } + published: + - { port: 443, on: home } # v4 reaches it only through the forward + inbound: allow # and v6 reaches it directly, so this matters + + workstation: # on the household network, not published + at: { segment: home, address: [192.168.1.250, 2001:db8:b:1::250] } + inbound: deny + + laptop: # starts at home; moves during the run + at: { segment: home, address: [192.168.1.98, 2001:db8:b:1::98] } + inbound: deny + +place: + all: [host] + anchor: [substrate] + +snapshot: raised +``` + +### What the ISP modem is doing here + +**Nothing, if it is in bridge mode** — which is the ordinary arrangement when the household has +its own router. A bridged modem is a media converter: it changes the physical medium and leaves +the packets alone, so it creates no IP-level fact and appears nowhere above. + +Were it in router mode it would be a second gateway, `home` would sit behind it rather than +behind `internet` directly, and publishing would need a rule on both — the case the model +cannot yet express. + +### What is deliberately absent + +Nothing here mentions overlay addresses, which node is the hub, who peers with whom, any name, +or any certificate. Research 004 recorded all of those for this topology, and **a scenario must +not state them** ([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)): they +are what the mesh does, and a scenario that supplied them would be certifying its own work. + +The absence is the point. Given the declaration above, whether a hub is elected, whether the +NATed machine's endpoint is learned, whether the roaming machine re-forms after moving — all of +it is observed. + +### Running it + +``` +move laptop → { segment: cafe, address: [10.50.3.23] } # it leaves the house +move laptop → detached # it sleeps +move laptop → { segment: home, address: [192.168.1.98, 2001:db8:b:1::98] } +``` + +One identity, four positions, one run. The `mapping_ttl: 30s` on `cafe` means a connection held +without refreshing dies while it is out — which is the point of putting a number there. + +Note that the laptop's three positions are on **three different public networks**: at home it +appears as `198.51.100.7`, at the café as `203.0.113.129`, and the machine it is trying to reach +is on a fourth. None of them are adjacent, all of them are routed. That is the situation the +mesh actually faces. + +## Is this general? — the axes a setup can vary along + +The question that matters is not *does this cover our mesh*, but **can it express any mesh**. + +The standard applied is not "every property a network has". It is **every property that changes +how the mesh behaves**. Bandwidth does not change correctness; MTU does. + +| Axis | Values | Expressible | | +|---|---|---|---| +| **Reachability** | attached · forwardable gateway · unforwardable gateway · isolated | yes | the core of the model, and **per family** | +| **Address family** | IPv4 · IPv6 · dual-stack · neither | yes | `cidr:` and `address:` take both; `nat:` names which families are translated | +| **Interfaces per machine** | one · several | yes | `at:` takes a list | +| **Reachability policy, host** | symmetric · asymmetric | yes | `inbound:` — a routable machine that refuses everything | +| **Reachability policy, network** | open · segmented · asymmetric between segments | yes | `policy:` — inter-segment rules, as a segmented router enforces | +| **Segments per gateway** | one · several behind one router | yes | identical gateway declarations mean one gateway machine | +| **Gateway state** | permanent · expiring mappings | yes | `mapping_ttl:` | +| **Path MTU** | standard · reduced | yes | `segments[].mtu` | +| **Overlapping ranges** | distinct · two sites both on `192.168.1.0/24` | yes | segments may carry the same range | +| **Segment count** | one · many · isolated island | yes | `kind:` distinguishes an island from a public network | +| **Adjacency** | same segment · routed · unrelated public networks | yes | public segments are routed, never bridged — so no two are adjacent unless declared so | +| **Gateway depth** | direct · one gateway · nested | **partly** | `to:` chains, so nesting exists; `published:` names one gateway, so forwarding through two does not | +| **Address stability** | static · dynamic · changes mid-run | **partly** | a machine can be *moved*; an address changing under it in place cannot be stated | +| **Path quality** | latency · loss · bandwidth | **no**, deliberately | changes performance, not correctness — modelling it makes a network simulator, not a fixture | + +### What closing the gaps changed + +**Address family was not a field.** It changed the position model: a machine behind a household +gateway is typically *unforwardable on v4 and directly attached on v6, simultaneously*. So +reachability is a property of *(machine, family)*, and *"can these two nodes reach each other"* +is no longer a yes/no question — it is asked once per family, and the asymmetric answers are the +interesting ones. That distinction did not exist in the model an hour ago and would have been +discovered by a mesh failing over one family while reporting the other. + +**`inbound:` became necessary because of v6.** With NAT, unreachability was implied by the +topology. With a globally routable v6 address, a machine is reachable unless something refuses — +so refusing has to be sayable, or v6 addressing would silently imply reachability. + +**`nat:` became a list rather than a boolean** for the same reason: a real gateway translates v4 +and routes v6, and a boolean cannot say that. + +### What remains open, and whether it matters + +Two partial axes, both extensible when something needs them, neither blocking: **nested +forwarding** and **an address changing in place**. A machine can already be moved, which covers +the roaming case; what is missing is a lease expiring underneath a machine that stays put. + +One deliberate exclusion: **path quality**. Latency and loss change how fast the mesh is, not +whether it is correct. If a timeout turns out to be load-bearing that judgement should be +revisited — and it would be revisited by a real failure, which is the right trigger. + +### The honest summary + +The model now covers **where a machine sits**, **what the path between machines is like**, and +**what is permitted between them** — across both address families. Together those are what the +mesh's reachability logic turns on. + +It contains almost no equipment, by the test above: a device that does not change what an IP +packet can do has no fault for a scenario to catch. What it does contain is every device that +does — which turns out to be the router, and a modem only when the modem is also a router. + +What it still does not model is *change over time* beyond moving a machine, *degradation* short +of failure, and *publishing through two gateways at once*. The first two are chosen; the third +is the one real absence, and it is exactly the double-NAT case. + ## Open - **`user` and `edge` profiles have no scenario.** A lab machine is always privileged, so the - two profiles that exist for unprivileged and phone-like participation cannot currently be - exercised. Either the lab grows a way to run the host unprivileged, or those profiles are - developed against something that is not a virtual machine. -- **Attaching and detaching during a run.** `segment: detached` covers a machine at rest; - moving one between segments while a scenario is live is what makes a roaming node - interesting, and that is lifecycle rather than declaration. + two profiles that exist for unprivileged and phone-like participation cannot be exercised. + Either the lab grows a way to run the host unprivileged, or those profiles are developed + against something that is not a virtual machine. This is the largest gap. - **Where `place:` gets its artifacts from.** Before the mesh is self-hosting these come from outside; afterwards from the mesh itself. The declaration should not have to care, which suggests a named source rather than a path. -- **Multiple scenarios at once.** Each needs its own segments and addresses. Whether the - declaration carries absolute addresses, as above, or a template the lab allocates from, - decides whether two scenarios can run side by side. +- **Nested forwarding** — `published:` names one gateway, so a machine behind two cannot be + published through both. +- **An address changing in place**, as a DHCP lease expiring under a machine that has not moved. diff --git a/03-DESIGN/01-to-be/03-scenario-lifecycle.md b/03-DESIGN/01-to-be/03-scenario-lifecycle.md new file mode 100644 index 0000000..8a06df0 --- /dev/null +++ b/03-DESIGN/01-to-be/03-scenario-lifecycle.md @@ -0,0 +1,140 @@ +--- +layer: to-be +status: designed +code: [mesh-lab] +updated: 2026-08-23 +decisions: + - 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md + - 02-DECISIONS/0031-the-lab-provides-the-underlay.md + - 02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md +--- + +# Scenario lifecycle + +The first thing the lab must do, and the only thing it must do before anything else can be +written: **materialise a mesh, return it to a known state, and destroy it** +([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)). + +A [declaration](02-scenario-declaration.md) describes a scenario. This describes what happens +to one. + +## The verbs + +| Verb | Does | +|---|---| +| `raise` | materialise a declaration into a running scenario instance | +| `snapshot` | name the current state of the whole scenario | +| `restore` | return the whole scenario to a named state | +| `move` | change a machine's position while the scenario runs | +| `exec` | run something on a machine, and get its output | +| `destroy` | tear the instance down | + +Six verbs, and `raise` plus `destroy` are enough to be useful. The rest are what make repetition +cheap, and cheap repetition is what turns the bootstrap path into an inner development loop +rather than a ceremony. + +## Raising, in order + +The order is not arbitrary — each step needs the one before it to exist: + +1. **Segments.** Isolated links, one per declared segment, belonging to this instance and + joined to nothing outside it + ([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)). +2. **Gateways.** Derived, never declared as machines: a gateway is materialised for each + distinct `gateway:` declaration, sitting on both its segment and its parent, carrying the + translation, forwarding and mapping-expiry the declaration asked for. +3. **Machines.** Each on its segments, holding its addresses. +4. **Policy.** Rules between segments, applied on the gateways that route between them. +5. **Placement.** Artifacts onto machines. +6. **Snapshot**, if the declaration named one. + +`raise` must wait for the machines to be **usable**, not for the call to return. Measured, those +are 14 seconds apart for a single machine — and reporting the earlier number would be the mesh's +own recurring failure, transport reported as effect. + +Raising is **convergent, not incremental**: raising an instance that already exists brings it to +the declared state rather than failing or duplicating. That is the same model the mesh itself +uses, and a lab that behaved differently from the thing it tests would be teaching the wrong +habit. + +## A failed raise leaves the wreckage + +A step that fails stops the raise +([ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md)) — and **does not tear +down**. + +Tearing down on failure destroys the only evidence of what went wrong, which is precisely +backwards: a scenario that failed to raise is more interesting than one that succeeded. The +instance stays, marked failed, with the step that failed named. + +The caller decides what happens next, and the two callers want different things +— the coordinator captures and destroys, a person opens a shell. That is the same +one-runner-two-callers split the lab design already makes, applied to failure. + +## Snapshots are whole-scenario + +A snapshot captures **every machine and the state of the network between them**, as one thing. +Restoring returns all of it. + +Per-machine snapshots would be cheaper and are wrong. The mesh keeps state that spans nodes — +what is assigned where, which grants exist, what has been delivered — so restoring one machine +to an earlier moment while its peers move on produces a mesh that has never existed and could +not. The faults found there would be artefacts of the lab. + +This is what makes *fresh* and *upgrade* both cheap and both default: one snapshot of a mesh +that has never seen a change, another of a mesh running the previous version, and a restore +between runs. + +## Moving a machine + +`move` changes a machine's position while the scenario runs: to another segment, with different +addresses, or to `detached`. + +It is the roaming case, and it is a **lifecycle** operation rather than a declaration because +the interesting part is the transition, not the destination. A mesh that forms correctly with a +node at home and correctly with it away may still fail to notice it moved. + +Moving does not invalidate a snapshot. A snapshot is a state to return to; a move is a change +made after it, and returning undoes it like any other change. + +## Reaching in + +Everything the lab does to a machine goes through the virtualisation layer, never over IP +([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)). `exec` runs a +command on a machine and returns its output. + +This has one consequence worth stating plainly: **a reachability question is asked from inside**. +*Can this machine reach that one* is `exec` on the first, testing the second — not a probe from +the workstation. The workstation is not on the scenario's network and its opinion of +reachability would be a different question with a misleadingly similar answer. + +Anything a person wants to open in a browser needs a deliberate forward out of the instance. +Nothing leaks by default. + +## The two callers + +The lab design already establishes that the runner serves the coordinator and a person, and +that anything only one of them can do will drift. Applied here: + +| | the coordinator | someone working on the mesh | +|---|---|---| +| wants | structured results, clean teardown, no prompts | readable output, the scenario **left standing**, one thing re-run without repeating the rest | +| on failure | capture, then destroy | leave it, open a shell | + +Both use the same verbs. The difference is what happens after the verdict, which is a caller's +decision rather than a second implementation. + +## Open + +- **What a snapshot costs — measured, and the answer is a blocker.** On the current host a + snapshot is a full copy: 1.6 GB and ten seconds for one small machine, over two minutes when + observed again, projecting to roughly a minute and a half for a four-machine reset-and-rerun + cycle. The cause is the storage driver, not virtual machines. See + [research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md). +- **Instance naming.** A declaration is a kind and instances are many; how they are named + decides whether a person can find the one they left standing yesterday. +- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying + the instance must not destroy them. +- **Placement before the mesh is self-hosting.** `place:` needs artifacts from somewhere, and + before the mesh builds itself that somewhere is outside it — the open question from + [research 006](../../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md). diff --git a/03-DESIGN/01-to-be/04-lab-installation.md b/03-DESIGN/01-to-be/04-lab-installation.md new file mode 100644 index 0000000..c1bd0cd --- /dev/null +++ b/03-DESIGN/01-to-be/04-lab-installation.md @@ -0,0 +1,116 @@ +--- +layer: to-be +status: designed +code: [mesh-lab] +updated: 2026-08-24 +decisions: + - 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md + - 02-DECISIONS/0008-a-failed-step-fails-the-job.md +--- + +# Installing the lab on a clean machine + +The lab has prerequisites — a virtualisation daemon, copy-on-write storage, a pool, an identity +permitted to talk to it — and it cannot get them from the mesh, because it is where the mesh is +built ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)). + +So the lab needs an install path of its own. This describes it, and the shape it has to take is +determined by two failures observed while measuring +([research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md)). + +## The two failures that shape this + +**One: installed is not available.** The virtualisation package was present and explicitly +installed. Both its units were disabled, the operator was in no group, and the client reported +the server unreachable. Nothing had failed — the declaration was satisfied exactly as written +([`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)). + +**Two, and worse: available is not adequate.** With the storage tooling absent, the daemon +offered one driver, everything worked, and snapshots took **seventy-six times longer** than they +needed to. Nothing failed. Nothing warned. A lab in that state runs correctly and is simply too +slow to use — and the inner loop it exists to provide quietly does not exist. + +The second is the more dangerous shape and it is a variant the mesh has not catalogued before. +Its usual failure is *reported success and did nothing*. This is **reported success and did it +seventy-six times slower**, which no error surface catches because nothing is wrong. + +## What follows: the lab verifies capability, never installation + +The install path may differ. **The verification does not.** + +Before the lab will raise anything, it asserts the outcomes it needs — not that packages are +present, but that the machine can actually do the work: + +| Assertion | Failing means | +|---|---| +| the daemon answers **as the invoking user**, not as root | a group membership that was granted but never took effect | +| a copy-on-write storage driver is offered | the userspace tooling is missing; snapshots will be full copies | +| **the pool the lab will use is on that driver** | a pool exists, and is the slow kind — the failure that has no symptom | +| hardware virtualisation is present | machines will be emulated and unusably slow | +| an image can be fetched or is cached | the first raise will fail late instead of early | + +Each check states **why it matters**, in the terms of what it costs. *"The pool uses the `dir` +driver"* means nothing to someone who does not already know it means seventy-six times slower +and unbounded at worst. + +**The lab refuses to run degraded.** It does not warn and continue: a warning about a slow inner +loop is read once and ignored forever, and the loop stays slow. This is +[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) applied where the failure is +performance rather than an error. + +## Two ways the prerequisites arrive + +**On a machine the mesh manages** — a module declares them, and a hook turns them into +capabilities: the units enabled, the group granted, the pool created on the right driver. This +already works; it is what was done for the virtualisation daemon itself. + +**On a machine the mesh does not manage** — the lab's own bootstrap does it. One command, on a +clean machine, that installs what is missing and configures it. + +The second path is not a convenience. It is **required**, because the lab must work before the +mesh does, and a lab that could only be installed by a mesh would be unable to host the +development of the mesh that installs it. + +Both paths end at the same verification. Whoever satisfied the prerequisites, the lab checks +them itself — because the lesson of the first failure is precisely that *something else said it +was done* is not evidence. + +## The lab is the second thing installed by hand + +Worth stating, because it looks like an exception and is not. + +The node host is the one thing installed by hand on a machine +([research 006](../../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)): everything else +arrives through it. The lab is the same shape on a workstation — installed once, by hand, and +then everything about the mesh is developed inside it. + +Two bootstraps, at two levels, for the same reason: **something has to be first, and pretending +otherwise produces a circularity that gets papered over with a script nobody exercises.** + +## What a clean install actually needs + +In order, on a machine with nothing: + +1. **A virtualisation daemon**, running, with its socket enabled. +2. **Copy-on-write storage tooling** — the kernel side is usually already present; it is the + userspace half that is missing and that decides whether the driver is offered at all. +3. **A pool on that driver.** A loop-backed file is sufficient and needs no partitioning, which + matters: requiring a dedicated filesystem would make the lab uninstallable on a machine + already in use. +4. **Group membership** for the operator — which does **not** apply to sessions that already + existed. Observed directly: a shell whose process tree predated the grant could not reach + the daemon while a fresh lookup showed the membership present. The bootstrap has to say so, + or the first thing a person meets is a permission error that looks like a broken install. +5. **Verification**, as above, before anything is raised. + +## Open + +- **Whether the lab's bootstrap may install packages at all**, given that the mesh's rules + forbid installing by hand. The resolution is probably that the lab's bootstrap *is* the + sanctioned mechanism on an unmanaged machine, in the way the mesh's own first-node script is — + but that is an argument to record, not to assume. +- **What "adequate" means numerically.** The checks above are qualitative. A snapshot-time + threshold would catch a copy-on-write pool that is slow for some other reason, and would be a + real assertion rather than a proxy. +- **Whether the lab should own its pool** rather than using an existing one. Owning it makes the + driver guaranteed; sharing it avoids duplicating storage on a machine that already has a pool. diff --git a/03-DESIGN/01-to-be/README.md b/03-DESIGN/01-to-be/README.md index e46578d..7f2fcc7 100644 --- a/03-DESIGN/01-to-be/README.md +++ b/03-DESIGN/01-to-be/README.md @@ -12,6 +12,8 @@ document is written and this one's status becomes `implemented`. | [`00-work-breakdown.md`](00-work-breakdown.md) | How the decomposition gets built, in what order, and where a human must look | [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) | | [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md), [0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) | | [`02-scenario-declaration.md`](02-scenario-declaration.md) | What a scenario declares — the underlay, and what to place on it | [ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md) | +| [`03-scenario-lifecycle.md`](03-scenario-lifecycle.md) | What happens to a scenario — raise, snapshot, restore, move, destroy | [ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md) | +| [`04-lab-installation.md`](04-lab-installation.md) | Getting the lab onto a clean machine, and why it verifies capability rather than installation | [ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) | ## Not yet written