From e65e5809dc541de36987e4c0a5bb6f4bb6209eeb Mon Sep 17 00:00:00 2001 From: jochen Date: Sun, 23 Aug 2026 22:42:27 +0200 Subject: [PATCH 1/9] The scenario declaration gets a real network model MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit forwarded: [443] was the tell. It implied a destination-NAT rule while never saying from which address, and the address is the whole point: a household's public address is what a peer records as the endpoint when a machine there dials out, and what a public name for a published machine there resolves to. It was decoration in the old shape and is load-bearing in this one. The model now names three positions a machine can be in, because they are genuinely different and the mesh has to cope with all three. Directly attached, with its own routable address. Behind a gateway you control, reachable only through a forwarded port at the gateway's address. Behind a gateway you do not control, reachable not at all, with an apparent address belonging to someone else's router that changes when the machine moves. The third is the hard one and the one that breaks reachability assumptions first. A gateway now carries three facts instead of a boolean: the parent segment, the address the world sees the network as, and whether addresses are translated — so a routed range is expressible as well as ordinary household NAT. published names the gateway it forwards through, which is how a machine on a LAN that itself has a public address is stated, and publishing on a foreign gateway is a declaration error because that is exactly the constraint being reproduced. Moving a machine between positions becomes a lifecycle operation rather than a declaration: the same identity at home, then on a foreign network, then asleep, in one run. Whether the overlay survives that and notices the endpoint changed is observed, never arranged. All three RFC 5737 ranges are now allocated a job — the internet segment, a foreign network, and a spare — with private segments kept byte-identical to production because those addresses mean the same everywhere. New open question worth having: a real gateway forgets NAT mappings after a timeout, and whether a scenario can say so decides whether keepalive behaviour is testable or merely hoped for. --- 03-DESIGN/01-to-be/02-scenario-declaration.md | 167 +++++++++++++----- 1 file changed, 122 insertions(+), 45 deletions(-) diff --git a/03-DESIGN/01-to-be/02-scenario-declaration.md b/03-DESIGN/01-to-be/02-scenario-declaration.md index a159a3a..84e40af 100644 --- a/03-DESIGN/01-to-be/02-scenario-declaration.md +++ b/03-DESIGN/01-to-be/02-scenario-declaration.md @@ -17,31 +17,60 @@ everything in the lab hangs off, so it is worth getting small. It states what a hosting provider and a home router would provide, and nothing the mesh is responsible for ([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)). +## Three positions a machine can be in + +The underlay's whole job is to reproduce **where a machine sits relative to the internet**, +because that is what the mesh has to cope with and what only production currently exercises. +There are three positions, and they are genuinely different: + +| Position | Reachable from outside | Address | Example | +|---|---|---|---| +| **Directly attached** | yes, at its own address | fixed, its own | a hosted server | +| **Behind a gateway you control** | only through a forwarded port, at the *gateway's* address | private, plus the gateway's public one | a machine at home | +| **Behind a gateway you don't control** | **no** | private, and it changes | a laptop on someone else's network | + +The third is the hard one and the reason this matters. A machine there can dial out and nothing +more: it cannot be published, its apparent address belongs to somebody else's router, and that +address changes when it moves. Every assumption a mesh makes about reachability breaks there +first. + +A declaration has to be able to say all three, and to move a machine between them. + ## The shape ```yaml -scenario: published-behind-nat +scenario: roaming-and-published segments: - wan: - cidr: 203.0.113.0/24 # RFC 5737 — never routes on the real internet - lan: + internet: + cidr: 203.0.113.0/24 # the simulated public internet, RFC 5737 + home: cidr: 192.168.1.0/24 - behind: wan # NAT; the lab materialises a router + gateway: + to: internet + address: 203.0.113.50 # what the world sees this network as + nat: true + elsewhere: # a network we do not control + cidr: 198.51.100.0/24 + gateway: + to: internet + address: 203.0.113.80 + nat: true machines: anchor: - segment: wan - address: 203.0.113.10 + at: { segment: internet, address: 203.0.113.10 } + home-server: - segment: lan - address: 192.168.1.135 - forwarded: [443] # reachable from wan through the router + at: { segment: home, address: 192.168.1.135 } + published: + - { port: 443, on: home } # DNAT: 203.0.113.50:443 → 192.168.1.135:443 + workstation: - segment: lan - address: 192.168.1.250 + at: { segment: home, address: 192.168.1.250 } + laptop: - segment: detached # reachable by nothing until attached + at: { segment: home, address: 192.168.1.98 } place: all: [host] @@ -50,41 +79,87 @@ place: snapshot: raised ``` -That is a complete bootstrap scenario. Nothing in it mentions the overlay, a hub, peering, -names or certificates — all of which are outcomes to be observed. +## What each part means, precisely -## The four parts +**`segments`** — a broadcast domain with an address range. A segment with no `gateway:` *is* +the internet as far as the scenario is concerned. A segment with one sits behind it. -**`segments`** — the networks that exist. `behind:` declares NAT, and is the only place a -router comes from: the lab materialises one without being asked, because NAT has to run -somewhere. This is the one implicit machine in an otherwise explicit declaration. +**`gateway:`** — how a segment reaches its parent, and this is where the previous version was +too thin. It carries three facts, and all three are load-bearing: -**`machines`** — what sits where. A machine has a segment and an address, and that is nearly -all. `forwarded:` opens a port through the router, which is what makes *published but behind -NAT* reproducible — the case that exists only in production today. `segment: detached` is a -machine on no network, which is how a roaming node is expressed at rest. +- `to:` — the parent segment. +- `address:` — **the address the outside world sees this network as.** For a household + connection this is the public address the ISP hands out. It is not decoration: it is what a + peer records as the endpoint when a machine here dials out, and what a public name for a + published machine here resolves to. +- `nat:` — whether addresses are translated. `true` gives the ordinary household case: many + private machines behind one public address. `false` describes a routed range, where machines + keep their own addresses and the gateway only forwards. -**`place`** — what goes inside. `all:` applies to every machine; a machine name overrides for -that machine. This is the only part that differs between the two scenario classes. +The lab materialises a machine to be the gateway. That is the one implicit machine in an +otherwise explicit declaration, and it exists because NAT has to run somewhere. -**`snapshot`** — names the state once placement finishes, so a run can return to it without -raising everything again. Snapshots are what make repetition cheap, and cheap repetition is -what makes the bootstrap path the inner development loop rather than a ceremony. +**`machines[].at`** — segment and address. That pair alone determines which of the three +positions a machine is in: on a gateway-less segment it is directly attached; on a segment with +a gateway it is behind one. + +**`machines[].published`** — a destination-NAT rule on a named gateway, stated as an outcome +rather than a port list. `{ port: 443, on: home }` means the `home` gateway forwards its own +`203.0.113.50:443` to this machine's `443`. The resulting public endpoint is derivable, which is +the point: a scenario never writes an endpoint down, and the mesh has to discover it. + +A machine may be published on **any gateway between it and the internet** — which is how *"our +LAN also has a public IP"* is expressed, and why `on:` names the gateway rather than being +implied. It cannot be published at all on a gateway the scenario models as foreign; attempting +it is a declaration error, because that is precisely the constraint being reproduced. + +**`at: detached`** — on no segment. A machine that exists and can reach nothing. + +## Moving a machine is a lifecycle operation + +`at:` states where a machine *starts*. Moving it is something a run does: + +``` +move laptop → { segment: elsewhere, address: 198.51.100.23 } +move laptop → detached +move laptop → { segment: home, address: 192.168.1.98 } +``` + +This is the roaming case made testable, and it is the one that finds the interesting faults. +The same machine, the same identity, three positions in one run: at home where its peers can +reach it directly, on a foreign network where it can only dial out and its apparent address +belongs to a router it does not control, and asleep. + +Whether the overlay survives that, re-forms, and is noticed to have changed endpoint is +**observed**, never arranged +([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)). ## Why the addresses are load-bearing -The routable segment uses RFC 5737 documentation space, and this is not a stylistic choice. +The internet segment uses RFC 5737 documentation space, and this is not a stylistic choice. The mesh decides *public versus private* by matching the address. A private range on the -segment meant to be routable makes the hub test as unreachable, and **the mesh silently never -forms** — no error, no failed step, just a mesh that does not exist. Research 004 calls this -the single most important fact in its analysis. +segment meant to be routable makes a would-be hub test as unreachable, and **the mesh silently +never forms** — no error, no failed step, just a mesh that does not exist. Research 004 calls +this the single most important fact in its analysis. -So the format should make this hard to get wrong rather than merely documented: a segment -without `behind:` is a routable segment, and an address in it that is not documentation space -is a declaration error, refused before anything is raised. That is +RFC 5737 reserves three ranges, which is exactly enough for the topology above: + +| Range | Used for | +|---|---| +| `203.0.113.0/24` | the internet segment itself — directly attached machines, and gateway addresses | +| `198.51.100.0/24` | a foreign network, so a roaming machine's apparent address is plainly not ours | +| `192.0.2.0/24` | spare — a second foreign network, or a second site | + +Private segments use RFC 1918 and can be **byte-identical to production**, because those +addresses mean the same thing everywhere. Only the public side is substituted, and only because +it must be. + +The format should make getting this wrong hard rather than merely documented: a segment without +a `gateway:` is a public segment, and an address in it — including a gateway's `address:` — that +is not documentation space is a declaration error, refused before anything is raised. That is [ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) applied to a configuration -file — the failure it prevents is silent, so the check has to be loud. +file: the failure it prevents is silent, so the check has to be loud. ## The same declaration serves both classes @@ -127,15 +202,17 @@ not first. ## Open - **`user` and `edge` profiles have no scenario.** A lab machine is always privileged, so the - two profiles that exist for unprivileged and phone-like participation cannot currently be - exercised. Either the lab grows a way to run the host unprivileged, or those profiles are - developed against something that is not a virtual machine. -- **Attaching and detaching during a run.** `segment: detached` covers a machine at rest; - moving one between segments while a scenario is live is what makes a roaming node - interesting, and that is lifecycle rather than declaration. + two profiles that exist for unprivileged and phone-like participation cannot be exercised. + Either the lab grows a way to run the host unprivileged, or those profiles are developed + against something that is not a virtual machine. This is the largest gap. - **Where `place:` gets its artifacts from.** Before the mesh is self-hosting these come from outside; afterwards from the mesh itself. The declaration should not have to care, which suggests a named source rather than a path. -- **Multiple scenarios at once.** Each needs its own segments and addresses. Whether the - declaration carries absolute addresses, as above, or a template the lab allocates from, - decides whether two scenarios can run side by side. +- **Multiple scenarios at once.** Each needs its own segments and addresses, and the shape above + writes addresses absolutely. Whether a scenario carries literal addresses or a template the + lab allocates from decides whether two can run side by side — and there are only three + documentation ranges to go round. +- **Gateway behaviour beyond forwarding.** A real household gateway also has a NAT table with + timeouts, and connection tracking that drops idle flows. Whether a scenario can express *"the + gateway forgets a mapping after N seconds"* decides whether keepalive behaviour is testable + or merely hoped for. -- 2.54.0 From b944904f1a975e4d3836682a2d5fa607ce6874dd Mon Sep 17 00:00:00 2001 From: jochen Date: Sun, 23 Aug 2026 22:48:46 +0200 Subject: [PATCH 2/9] Audit the scenario model for generality, and fix what it found MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The question is not whether the model covers our mesh but whether it can express any mesh. Audited against the axes a deployment varies along, with the standard being every property that changes how the mesh BEHAVES rather than every property a network has — bandwidth does not change correctness, MTU does. One real bug, now fixed. A segment with no gateway was read as the internet, which made an isolated network inexpressible: a LAN with no route out would have been treated as public and forced onto documentation addresses. Segments now state kind: public or private, and a private segment with no gateway is an island. A mesh spanning a site with no internet is a real topology. One modelling error, now corrected. The three positions were framed by ownership — a gateway you control versus one you do not. The axis is forwardability. Carrier-grade NAT is your own connection and is still unforwardable, so it belongs with the café network. Gateways gain forwardable:, independent of nat:, and publishing through an unforwardable one is a declaration error because that is the constraint being reproduced. Three genuine gaps recorded in priority order. Address family: cidr is implicitly v4, and a v6-only node is not exotic — a mesh that assumes v4 fails there completely rather than partially, which makes this a second world rather than a refinement. Expiring NAT mappings: without them keepalive behaviour is hoped for rather than tested, and for a mesh mostly behind NAT that is the fault that shows up after an idle night. MTU: tunnels fragment, and a smaller-MTU path establishes a connection that then silently drops large packets — the exact shape this effort exists to stop shipping. Latency and loss are deliberately out: they change performance, not correctness, and modelling them makes a network simulator rather than a fixture. Also adds a NAT primer, because the three positions are consequences of it and the document should not assume the reader already knows why a mesh dials outward and never inward. --- 03-DESIGN/01-to-be/02-scenario-declaration.md | 133 ++++++++++++++++-- 1 file changed, 118 insertions(+), 15 deletions(-) diff --git a/03-DESIGN/01-to-be/02-scenario-declaration.md b/03-DESIGN/01-to-be/02-scenario-declaration.md index 84e40af..38f6fd0 100644 --- a/03-DESIGN/01-to-be/02-scenario-declaration.md +++ b/03-DESIGN/01-to-be/02-scenario-declaration.md @@ -17,22 +17,57 @@ everything in the lab hangs off, so it is worth getting small. It states what a hosting provider and a home router would provide, and nothing the mesh is responsible for ([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)). +## What NAT does, and why the design turns on it + +A household or office has **one** address the outside world can see, and **many** machines +behind it. Network address translation is what reconciles those. + +When a machine inside dials out, the gateway rewrites the packet's source from the private +address to the public one, **remembers the mapping**, and rewrites the replies on the way back. +Four consequences follow, and every one of them shapes this design: + +1. **Outbound works; inbound does not.** A mapping exists only because something inside started + a conversation. Nothing outside can start one — there is no mapping to look up, and no way + to know which internal machine was meant. +2. **A forwarded port is a permanent mapping made by hand**, in the inbound direction: + *anything arriving at the public address on 443 goes to this machine.* That is the only way + a machine behind NAT becomes reachable, and it requires control of the gateway. +3. **Mappings expire.** A gateway forgets one that goes unused. This is why anything holding a + connection through NAT sends keepalives, and why a mesh that does not is fine until it is + idle. +4. **From outside, every machine behind the gateway looks like one address.** Identity and + address stop corresponding. + +This is why the mesh dials outward and never inward +([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md)), why a hub exists at +all, and why a node's endpoint is something a peer **learns** from arriving packets rather than +something anyone configures. + +**Carrier-grade NAT** is the same mechanism applied by an ISP: your own gateway gets a private +address too, and the public one is shared with strangers. Nothing can be forwarded, because the +rule would have to live on equipment you do not own. Common on mobile connections and +increasingly on fixed ones. + ## Three positions a machine can be in The underlay's whole job is to reproduce **where a machine sits relative to the internet**, because that is what the mesh has to cope with and what only production currently exercises. There are three positions, and they are genuinely different: -| Position | Reachable from outside | Address | Example | +| Position | Reachable from outside | Apparent address | Example | |---|---|---|---| -| **Directly attached** | yes, at its own address | fixed, its own | a hosted server | -| **Behind a gateway you control** | only through a forwarded port, at the *gateway's* address | private, plus the gateway's public one | a machine at home | -| **Behind a gateway you don't control** | **no** | private, and it changes | a laptop on someone else's network | +| **Attached** | yes, at its own address | its own | a hosted server | +| **Behind a forwardable gateway** | only through a forwarded port, at the *gateway's* address | the gateway's | a machine at home | +| **Behind an unforwardable gateway** | **no** | someone else's, and it changes | a laptop on a café network; anything behind carrier-grade NAT | -The third is the hard one and the reason this matters. A machine there can dial out and nothing -more: it cannot be published, its apparent address belongs to somebody else's router, and that -address changes when it moves. Every assumption a mesh makes about reachability breaks there -first. +The axis is **forwardability, not ownership** — which is worth stating because the obvious +framing gets it wrong. Carrier-grade NAT is *your* connection and is still unforwardable, so it +belongs in the third row alongside the café. What the mesh has to cope with is whether an +inbound mapping can be made, not who owns the equipment. + +The third position is the hard one. A machine there can dial out and nothing more: it cannot be +published, its apparent address belongs to a router it does not control, and that address +changes when it moves. Every assumption a mesh makes about reachability breaks there first. A declaration has to be able to say all three, and to move a machine between them. @@ -43,19 +78,24 @@ scenario: roaming-and-published segments: internet: - cidr: 203.0.113.0/24 # the simulated public internet, RFC 5737 + kind: public # stands in for the internet — RFC 5737 addresses + cidr: 203.0.113.0/24 home: + kind: private cidr: 192.168.1.0/24 gateway: to: internet address: 203.0.113.50 # what the world sees this network as nat: true + forwardable: true # we control it, so ports can be opened elsewhere: # a network we do not control + kind: private cidr: 198.51.100.0/24 gateway: to: internet address: 203.0.113.80 nat: true + forwardable: false # café wifi, or carrier-grade NAT machines: anchor: @@ -81,8 +121,14 @@ snapshot: raised ## What each part means, precisely -**`segments`** — a broadcast domain with an address range. A segment with no `gateway:` *is* -the internet as far as the scenario is concerned. A segment with one sits behind it. +**`segments`** — a broadcast domain with an address range, and a `kind:`. + +`kind: public` marks the segment that stands in for the internet. `kind: private` is everything +else. This is stated rather than inferred, and the earlier version inferred it — *a segment with +no gateway is the internet* — which made an **isolated network inexpressible**: a LAN with no +route out is a private segment with no gateway, and would have been read as the internet and +forced to use documentation addresses. A mesh spanning a site with no internet access is a real +topology, and the model has to be able to say it. **`gateway:`** — how a segment reaches its parent, and this is where the previous version was too thin. It carries three facts, and all three are load-bearing: @@ -95,6 +141,10 @@ too thin. It carries three facts, and all three are load-bearing: - `nat:` — whether addresses are translated. `true` gives the ordinary household case: many private machines behind one public address. `false` describes a routed range, where machines keep their own addresses and the gateway only forwards. +- `forwardable:` — whether an inbound mapping can be created. Independent of `nat:`, and the + field that separates a home gateway from carrier-grade NAT. Publishing through a gateway with + `forwardable: false` is a declaration error, because that is exactly the constraint being + reproduced. The lab materialises a machine to be the gateway. That is the one implicit machine in an otherwise explicit declaration, and it exists because NAT has to run somewhere. @@ -199,6 +249,61 @@ not first. - **Steps.** A scenario is a desired state. Anything expressed as an ordered list of actions belongs in the lifecycle, not the declaration. +## Is this general? — the axes a setup can vary along + +The question that matters is not *does this cover our mesh*, but **can it express any mesh**. +Audited against the axes a real deployment varies along, the answer is *most, deliberately not +all, and three genuine gaps*. + +The standard applied is not "every property a network has". It is **every property that changes +how the mesh behaves**. Bandwidth does not change correctness; MTU does. + +| Axis | Values | Expressible | | +|---|---|---|---| +| **Reachability** | attached · forwardable gateway · unforwardable gateway · isolated | yes | the core of the model | +| **Address stability** | static · dynamic · changes mid-run | **partly** | a machine can be *moved*, but an address that changes under it cannot be stated | +| **Gateway depth** | direct · one gateway · nested gateways | **partly** | `to:` chains, so nesting exists; `published:` names one gateway, so forwarding through two does not | +| **Address family** | IPv4 · IPv6 · dual-stack | **no** | `cidr:` is implicitly v4. A v6-only node is a real topology and cannot be written | +| **Interfaces per machine** | one · several | **no** | `at:` is singular. A multi-homed node — on a LAN and a WAN at once — is inexpressible | +| **Path properties** | MTU · latency · loss | **no** | MTU matters: tunnels fragment, and a lower-MTU path is a classic silent failure | +| **Reachability policy** | symmetric · asymmetric | **no** | a firewall dropping inbound while outbound works is different from NAT and behaves differently | +| **Gateway state** | permanent · expiring mappings | **no** | mappings time out; whether keepalives work is untestable without it | +| **Overlapping ranges** | distinct · two sites both on `192.168.1.0/24` | yes | two segments may carry the same range — common, and it breaks routing | +| **Segment count** | one · many · isolated island | yes | after the `kind:` fix above | + +### What this says + +**Three gaps are real and should be closed**, in this order: + +1. **Address family.** A v6-only or dual-stack node is not exotic, and a mesh that assumes v4 + fails there completely rather than partially. This is the largest gap. +2. **Expiring NAT mappings.** Without it, keepalive behaviour is hoped for rather than tested — + and for a mesh where most nodes sit behind NAT, that is the failure mode most likely to + appear only after everything has been idle overnight. +3. **MTU.** Tunnels fragment. A path with a smaller MTU produces a connection that establishes + and then silently drops large packets, which is exactly the shape of fault this whole effort + exists to stop shipping. + +**Two are deliberately out of scope** unless something argues otherwise: latency and loss. +They change performance, not correctness, and a scenario that models them is a network +simulator rather than a fixture. + +**Two are partial and probably fine for now:** nested forwarding and mid-run address change. +Both are expressible with small extensions when something needs them, and neither blocks the +bootstrap scenario. + +### The honest summary + +The model covers **where a machine sits**, which is what the mesh's reachability logic turns +on, and it now covers it completely. It does not yet cover **what the path between machines is +like**, and one of those — address family — is not a refinement but a second world the mesh +would have to work in. + +None of this blocks phase 0. A bootstrap scenario is one machine and a pinned bundle, and needs +none of it. But the gaps should be closed before the lab is trusted to say a mesh *works*, +because today it could only say it works over IPv4, on an unconstrained path, against gateways +that never forget. + ## Open - **`user` and `edge` profiles have no scenario.** A lab machine is always privileged, so the @@ -212,7 +317,5 @@ not first. writes addresses absolutely. Whether a scenario carries literal addresses or a template the lab allocates from decides whether two can run side by side — and there are only three documentation ranges to go round. -- **Gateway behaviour beyond forwarding.** A real household gateway also has a NAT table with - timeouts, and connection tracking that drops idle flows. Whether a scenario can express *"the - gateway forgets a mapping after N seconds"* decides whether keepalive behaviour is testable - or merely hoped for. +- **The three gaps from the audit above** — address family, expiring NAT mappings, MTU — in + that order. The first is the one that is a second world rather than a refinement. -- 2.54.0 From 274bd3b304653398177d29b47500f0cb31f028ed Mon Sep 17 00:00:00 2001 From: jochen Date: Sun, 23 Aug 2026 22:59:31 +0200 Subject: [PATCH 3/9] =?UTF-8?q?Close=20the=20missing=20axes=20=E2=80=94=20?= =?UTF-8?q?and=20address=20family=20changes=20the=20model?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Address family was not a field. IPv6 usually has no NAT, so a machine behind a household gateway is typically unforwardable on v4 and DIRECTLY ATTACHED on v6, at the same moment. The three positions therefore apply per family, and reachability is a property of (machine, family) rather than of a machine. The consequence is bigger than the syntax: 'can these two nodes reach each other' stops being a yes/no question. It is asked once per family, and the asymmetric answers are the interesting ones. A mesh treating reachability as one fact per node reaches a peer over one family, fails over the other, and reports whichever it tried. That distinction did not exist in the model and would have been found by a failure rather than by reading. Two fields follow from it. inbound: allow|deny became necessary because with NAT unreachability was implied by topology, while a globally routable v6 address is reachable unless something refuses — so refusing has to be sayable or v6 addressing silently implies reachability. And nat: became a list of families rather than a boolean, because a real gateway translates v4 and routes v6 and a boolean cannot say that. mapping_ttl closes the keepalive gap: a mesh holding a connection through NAT without refreshing it works perfectly until the far side goes quiet for longer than the mapping lives. segments[].mtu closes the fragmentation gap: an overlay adds a header, so a tunnel over a reduced-MTU path establishes a connection and then silently drops large packets. at: takes a list, so a multi-homed machine is expressible — which the model already implicitly required, since a border machine sits on two segments. v6 uses RFC 3849 documentation space, the exact counterpart of the RFC 5737 rule and load-bearing for the same reason. Remaining: nested forwarding and an address changing in place, both extensible when needed. Path quality stays deliberately out — it changes performance, not correctness, and modelling it makes a network simulator rather than a fixture. --- 03-DESIGN/01-to-be/02-scenario-declaration.md | 174 ++++++++++++------ 1 file changed, 117 insertions(+), 57 deletions(-) diff --git a/03-DESIGN/01-to-be/02-scenario-declaration.md b/03-DESIGN/01-to-be/02-scenario-declaration.md index 38f6fd0..a818906 100644 --- a/03-DESIGN/01-to-be/02-scenario-declaration.md +++ b/03-DESIGN/01-to-be/02-scenario-declaration.md @@ -71,6 +71,32 @@ changes when it moves. Every assumption a mesh makes about reachability breaks t A declaration has to be able to say all three, and to move a machine between them. +## Reachability is per address family, not per machine + +Adding IPv6 is not a field. It changes the position model, and the reason is worth stating +before the syntax. + +**IPv6 usually has no NAT.** A machine behind a household gateway can hold a *globally routable* +v6 address while its v4 address is private and unforwardable. The same machine, at the same +moment, is in **two different positions at once**: + +| | IPv4 | IPv6 | +|---|---|---| +| a typical machine at home | behind an unforwardable-or-forwardable gateway | **attached**, directly reachable | +| a machine on mobile data | behind carrier-grade NAT | often attached, sometimes absent entirely | +| a machine on an older network | attached or behind NAT | **no address at all** | + +So the three positions apply **per family**, and a machine's reachability is a property of +*(machine, family)* rather than of the machine. A mesh that treats reachability as one fact per +node will reach a peer over one family, fail over the other, and report whichever it tried. + +That has a direct consequence for what the lab is for: *"can these two nodes reach each other"* +stops being a yes/no question. It is asked once per family, and the interesting answers are the +asymmetric ones. + +The v6 documentation prefix is `2001:db8::/32` (RFC 3849) — the exact counterpart of the RFC +5737 rule, and load-bearing for the same reason. + ## The shape ```yaml @@ -78,39 +104,52 @@ scenario: roaming-and-published segments: internet: - kind: public # stands in for the internet — RFC 5737 addresses - cidr: 203.0.113.0/24 + kind: public # RFC 5737 for v4, RFC 3849 for v6 + cidr: [203.0.113.0/24, 2001:db8::/32] + home: kind: private - cidr: 192.168.1.0/24 + cidr: [192.168.1.0/24, 2001:db8:1::/64] + mtu: 1500 gateway: to: internet - address: 203.0.113.50 # what the world sees this network as - nat: true - forwardable: true # we control it, so ports can be opened + address: 203.0.113.50 # what the world sees this network as, on v4 + nat: [v4] # v4 is translated; v6 is routed, not translated + forwardable: true + mapping_ttl: 120s # an unused inbound mapping is forgotten after this + elsewhere: # a network we do not control kind: private - cidr: 198.51.100.0/24 + cidr: [198.51.100.0/24] # v4 only — no v6 offered here at all + mtu: 1400 # a tunnelled path, smaller than standard gateway: to: internet address: 203.0.113.80 - nat: true + nat: [v4] forwardable: false # café wifi, or carrier-grade NAT + mapping_ttl: 30s # aggressive, as carrier NAT tends to be machines: anchor: - at: { segment: internet, address: 203.0.113.10 } + at: { segment: internet, address: [203.0.113.10, 2001:db8::10] } home-server: - at: { segment: home, address: 192.168.1.135 } + at: { segment: home, address: [192.168.1.135, 2001:db8:1::135] } published: - - { port: 443, on: home } # DNAT: 203.0.113.50:443 → 192.168.1.135:443 + - { port: 443, on: home } # v4 only: 203.0.113.50:443 → 192.168.1.135:443 + inbound: allow # v6 is routable here, so this decides whether it is reachable workstation: - at: { segment: home, address: 192.168.1.250 } + at: { segment: home, address: [192.168.1.250, 2001:db8:1::250] } + inbound: deny # a host firewall: dials out, accepts nothing laptop: - at: { segment: home, address: 192.168.1.98 } + at: { segment: home, address: [192.168.1.98, 2001:db8:1::98] } + + border: # a machine on two segments at once + at: + - { segment: home, address: [192.168.1.2] } + - { segment: internet, address: [203.0.113.60] } place: all: [host] @@ -138,20 +177,40 @@ too thin. It carries three facts, and all three are load-bearing: connection this is the public address the ISP hands out. It is not decoration: it is what a peer records as the endpoint when a machine here dials out, and what a public name for a published machine here resolves to. -- `nat:` — whether addresses are translated. `true` gives the ordinary household case: many - private machines behind one public address. `false` describes a routed range, where machines - keep their own addresses and the gateway only forwards. +- `nat:` — **which families are translated**, as a list. `[v4]` is the ordinary modern case: + v4 translated, v6 routed. `[v4, v6]` describes a gateway that translates both, which exists + and is worth being able to reproduce. `[]` is a routed range, where machines keep their own + addresses and the gateway only forwards. - `forwardable:` — whether an inbound mapping can be created. Independent of `nat:`, and the field that separates a home gateway from carrier-grade NAT. Publishing through a gateway with `forwardable: false` is a declaration error, because that is exactly the constraint being reproduced. +- `mapping_ttl:` — how long an unused inbound mapping survives. This is what makes keepalive + behaviour testable: a mesh that holds a connection through NAT without refreshing it works + perfectly until the far side goes quiet for longer than this. Aggressive values reproduce + carrier NAT; omitting it means mappings never expire, which no real gateway does. + +**`segments[].mtu`** — the largest packet the segment carries, defaulting to 1500. Lower values +reproduce tunnelled and PPPoE paths. This matters because an overlay adds its own header: a +tunnel over a 1400-byte path establishes a connection and then silently drops large packets, +which is the shape of fault this whole effort exists to stop shipping. + +**`machines[].inbound`** — `allow` or `deny`, a host firewall. Distinct from NAT and behaves +differently: a machine can be perfectly routable and still refuse everything unsolicited, which +is the normal state of a v6-addressed machine. Without this, v6 addressing would imply +reachability, and it does not. The lab materialises a machine to be the gateway. That is the one implicit machine in an otherwise explicit declaration, and it exists because NAT has to run somewhere. -**`machines[].at`** — segment and address. That pair alone determines which of the three -positions a machine is in: on a gateway-less segment it is directly attached; on a segment with -a gateway it is behind one. +**`machines[].at`** — segment and addresses, or a **list** of them for a machine on several +segments at once. Multi-homing is not exotic: it is what a border machine is, and what any node +with both a LAN and a WAN interface is. Each entry carries the addresses that machine holds on +that segment, one per family. + +Position follows from the pair, per family: on a `kind: public` segment a machine is attached; +on a private one it is behind that segment's gateway, unless the gateway does not translate +that family — in which case it is attached on that family and behind a gateway on the other. **`machines[].published`** — a destination-NAT rule on a named gateway, stated as an outcome rather than a port list. `{ port: 443, on: home }` means the `home` gateway forwards its own @@ -252,57 +311,57 @@ not first. ## Is this general? — the axes a setup can vary along The question that matters is not *does this cover our mesh*, but **can it express any mesh**. -Audited against the axes a real deployment varies along, the answer is *most, deliberately not -all, and three genuine gaps*. The standard applied is not "every property a network has". It is **every property that changes how the mesh behaves**. Bandwidth does not change correctness; MTU does. | Axis | Values | Expressible | | |---|---|---|---| -| **Reachability** | attached · forwardable gateway · unforwardable gateway · isolated | yes | the core of the model | -| **Address stability** | static · dynamic · changes mid-run | **partly** | a machine can be *moved*, but an address that changes under it cannot be stated | -| **Gateway depth** | direct · one gateway · nested gateways | **partly** | `to:` chains, so nesting exists; `published:` names one gateway, so forwarding through two does not | -| **Address family** | IPv4 · IPv6 · dual-stack | **no** | `cidr:` is implicitly v4. A v6-only node is a real topology and cannot be written | -| **Interfaces per machine** | one · several | **no** | `at:` is singular. A multi-homed node — on a LAN and a WAN at once — is inexpressible | -| **Path properties** | MTU · latency · loss | **no** | MTU matters: tunnels fragment, and a lower-MTU path is a classic silent failure | -| **Reachability policy** | symmetric · asymmetric | **no** | a firewall dropping inbound while outbound works is different from NAT and behaves differently | -| **Gateway state** | permanent · expiring mappings | **no** | mappings time out; whether keepalives work is untestable without it | -| **Overlapping ranges** | distinct · two sites both on `192.168.1.0/24` | yes | two segments may carry the same range — common, and it breaks routing | -| **Segment count** | one · many · isolated island | yes | after the `kind:` fix above | +| **Reachability** | attached · forwardable gateway · unforwardable gateway · isolated | yes | the core of the model, and **per family** | +| **Address family** | IPv4 · IPv6 · dual-stack · neither | yes | `cidr:` and `address:` take both; `nat:` names which families are translated | +| **Interfaces per machine** | one · several | yes | `at:` takes a list | +| **Reachability policy** | symmetric · asymmetric | yes | `inbound:` — a routable machine that refuses everything | +| **Gateway state** | permanent · expiring mappings | yes | `mapping_ttl:` | +| **Path MTU** | standard · reduced | yes | `segments[].mtu` | +| **Overlapping ranges** | distinct · two sites both on `192.168.1.0/24` | yes | segments may carry the same range | +| **Segment count** | one · many · isolated island | yes | `kind:` distinguishes an island from the internet | +| **Gateway depth** | direct · one gateway · nested | **partly** | `to:` chains, so nesting exists; `published:` names one gateway, so forwarding through two does not | +| **Address stability** | static · dynamic · changes mid-run | **partly** | a machine can be *moved*; an address changing under it in place cannot be stated | +| **Path quality** | latency · loss · bandwidth | **no**, deliberately | changes performance, not correctness — modelling it makes a network simulator, not a fixture | -### What this says +### What closing the gaps changed -**Three gaps are real and should be closed**, in this order: +**Address family was not a field.** It changed the position model: a machine behind a household +gateway is typically *unforwardable on v4 and directly attached on v6, simultaneously*. So +reachability is a property of *(machine, family)*, and *"can these two nodes reach each other"* +is no longer a yes/no question — it is asked once per family, and the asymmetric answers are the +interesting ones. That distinction did not exist in the model an hour ago and would have been +discovered by a mesh failing over one family while reporting the other. -1. **Address family.** A v6-only or dual-stack node is not exotic, and a mesh that assumes v4 - fails there completely rather than partially. This is the largest gap. -2. **Expiring NAT mappings.** Without it, keepalive behaviour is hoped for rather than tested — - and for a mesh where most nodes sit behind NAT, that is the failure mode most likely to - appear only after everything has been idle overnight. -3. **MTU.** Tunnels fragment. A path with a smaller MTU produces a connection that establishes - and then silently drops large packets, which is exactly the shape of fault this whole effort - exists to stop shipping. +**`inbound:` became necessary because of v6.** With NAT, unreachability was implied by the +topology. With a globally routable v6 address, a machine is reachable unless something refuses — +so refusing has to be sayable, or v6 addressing would silently imply reachability. -**Two are deliberately out of scope** unless something argues otherwise: latency and loss. -They change performance, not correctness, and a scenario that models them is a network -simulator rather than a fixture. +**`nat:` became a list rather than a boolean** for the same reason: a real gateway translates v4 +and routes v6, and a boolean cannot say that. -**Two are partial and probably fine for now:** nested forwarding and mid-run address change. -Both are expressible with small extensions when something needs them, and neither blocks the -bootstrap scenario. +### What remains open, and whether it matters + +Two partial axes, both extensible when something needs them, neither blocking: **nested +forwarding** and **an address changing in place**. A machine can already be moved, which covers +the roaming case; what is missing is a lease expiring underneath a machine that stays put. + +One deliberate exclusion: **path quality**. Latency and loss change how fast the mesh is, not +whether it is correct. If a timeout turns out to be load-bearing that judgement should be +revisited — and it would be revisited by a real failure, which is the right trigger. ### The honest summary -The model covers **where a machine sits**, which is what the mesh's reachability logic turns -on, and it now covers it completely. It does not yet cover **what the path between machines is -like**, and one of those — address family — is not a refinement but a second world the mesh -would have to work in. +The model now covers **where a machine sits** and **what the path between machines is like**, +across both address families, which together are what the mesh's reachability logic turns on. -None of this blocks phase 0. A bootstrap scenario is one machine and a pinned bundle, and needs -none of it. But the gaps should be closed before the lab is trusted to say a mesh *works*, -because today it could only say it works over IPv4, on an unconstrained path, against gateways -that never forget. +What it does not model is *change over time* beyond moving a machine, and *degradation* short of +failure. Both are absences chosen rather than overlooked. ## Open @@ -317,5 +376,6 @@ that never forget. writes addresses absolutely. Whether a scenario carries literal addresses or a template the lab allocates from decides whether two can run side by side — and there are only three documentation ranges to go round. -- **The three gaps from the audit above** — address family, expiring NAT mappings, MTU — in - that order. The first is the one that is a second world rather than a refinement. +- **Nested forwarding** — `published:` names one gateway, so a machine behind two cannot be + published through both. +- **An address changing in place**, as a DHCP lease expiring under a machine that has not moved. -- 2.54.0 From b1874f1d0fbfc57c9b4b71adaa20764977f5bf3a Mon Sep 17 00:00:00 2001 From: jochen Date: Sun, 23 Aug 2026 23:32:30 +0200 Subject: [PATCH 4/9] Segment policy, shared gateways, and what the model leaves out MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Asked whether a real setup is coverable — router, modem, access points — the answer splits, and one part was a genuine gap. Most equipment is invisible and the omission is deliberate. The test: does the device change what an IP packet can do? A switch moves frames within a segment. An access point bridges wireless clients onto one — a machine on wifi and a machine on cable are the same machine to IP. A controller configures equipment and has no packets of its own. Modelling any of them adds a fixture with no fault to catch. Two entries in that list do matter. A modem in bridge mode is a media converter and invisible; in router mode it is a second gateway, which is double NAT — expressible as nested segments, but publishing through two gateways still is not, and that is now named as the one real absence. And VLANs are segments, which exposed the gap: inter-segment policy was inexpressible. inbound: is a HOST firewall, per machine. A segmented router enforcing rules between networks is a different thing and blocks traffic regardless of what the destination thinks — a node behind such a rule cannot be reached even by a peer that knows exactly where it is. policy: states it as a fact about a pair rather than a property of either, defaulting to allowed and asymmetric by design, because the useful configuration is almost always one-directional. Segments may also share a gateway: identical gateway declarations mean one gateway machine, not two, because that is what a VLAN-capable router is — and two routers sharing an address would not work anyway. --- 03-DESIGN/01-to-be/02-scenario-declaration.md | 98 ++++++++++++++++++- 1 file changed, 93 insertions(+), 5 deletions(-) diff --git a/03-DESIGN/01-to-be/02-scenario-declaration.md b/03-DESIGN/01-to-be/02-scenario-declaration.md index a818906..62e59b2 100644 --- a/03-DESIGN/01-to-be/02-scenario-declaration.md +++ b/03-DESIGN/01-to-be/02-scenario-declaration.md @@ -308,6 +308,86 @@ not first. - **Steps.** A scenario is a desired state. Anything expressed as an ordered list of actions belongs in the lifecycle, not the declaration. +## What the model deliberately does not contain + +A real network is full of equipment: a modem, a router, switches, access points, controllers. +Almost none of it appears here, and the omission is deliberate rather than an oversight. + +**The test is whether a device changes what an IP packet can do.** If two machines can exchange +packets, at the same addresses, with the same reachability and the same MTU, whether or not the +device exists — then the device is invisible to the mesh, and modelling it would add a fixture +with no fault to catch. + +| Equipment | Modelled? | Why | +|---|---|---| +| **Switch** | no | Moves frames within a segment. Two machines on a switch are two machines on a segment. | +| **Access point** | no | Bridges wireless clients onto a segment. A machine on wifi and a machine on cable are the same machine to IP. | +| **Network controller** | no | Configures equipment. Its effects appear as segments and policy; it has no packets of its own. | +| **Cabling, PoE, uplink speed** | no | Change performance and availability, not reachability. | +| **Router / security gateway** | **yes — it *is* the gateway** | Translation, forwarding and inter-segment policy all live here. | +| **ISP modem** | **only in router mode** | In bridge mode it is a media converter and invisible. Doing its own NAT, it is a second gateway — and that is double NAT. | +| **VLANs** | **yes — they are segments** | Machines on separate VLANs cannot reach each other except through the router, which is the definition of a separate segment. | +| **Inter-VLAN firewall rules** | **yes** — see below | A rule stopping one segment reaching another is a reachability fact, and the mesh will hit it. | + +The two entries worth dwelling on are the ones where a common household setup produces +something the mesh has to survive. + +**A modem in router mode gives you double NAT.** Your gateway holds a private address from the +modem, which holds the public one. Forwarding then requires a rule on *both*, and one of them +may not be configurable. This is expressible as nested segments — `to:` chains — but publishing +through two gateways is not, and remains open. + +**Segmented networks are the common case, not the exotic one.** A router with separate networks +for trusted machines, guests and devices is ordinary, and the rules between them are ordinary +too. A mesh node on one segment and a mesh node on another are, as far as reachability goes, on +different networks that happen to share a gateway. + +## Segments may share a gateway + +Several segments can name the same parent and the same external address. That is one router +with several networks behind it, which is what a VLAN-capable gateway is: + +```yaml +segments: + trusted: + kind: private + cidr: [192.168.1.0/24] + gateway: { to: internet, address: 203.0.113.50, nat: [v4], forwardable: true } + devices: + kind: private + cidr: [192.168.30.0/24] + gateway: { to: internet, address: 203.0.113.50, nat: [v4], forwardable: true } +``` + +Identical gateway declarations mean **one gateway machine**, not two. The lab materialises a +single router serving both, because that is what the topology being reproduced is — and two +routers sharing one address would not work anyway. + +## Policy between segments + +Sharing a gateway does not mean segments can reach each other. What they may do is stated +separately, because it is a fact about a pair rather than a property of either: + +```yaml +policy: + - { from: devices, to: trusted, allow: false } # devices may not initiate to trusted + - { from: trusted, to: devices, allow: true } # the reverse is fine +``` + +Default is `true` between segments behind the same gateway, matching a router with no rules +configured. Asymmetry is the normal case and the reason this cannot be a single flag: the +useful configuration is almost always one-directional. + +This is **not** the same as `machines[].inbound`, and conflating them loses a real distinction: + +| | Enforced by | Blocks | +|---|---|---| +| `policy` | the gateway, between segments | everything crossing, regardless of what the destination thinks | +| `inbound` | the machine itself | unsolicited traffic that already reached it | + +A mesh node behind a `policy` deny cannot be reached even by a peer that knows exactly where it +is — and that is a real topology, not a contrived one. + ## Is this general? — the axes a setup can vary along The question that matters is not *does this cover our mesh*, but **can it express any mesh**. @@ -320,7 +400,9 @@ how the mesh behaves**. Bandwidth does not change correctness; MTU does. | **Reachability** | attached · forwardable gateway · unforwardable gateway · isolated | yes | the core of the model, and **per family** | | **Address family** | IPv4 · IPv6 · dual-stack · neither | yes | `cidr:` and `address:` take both; `nat:` names which families are translated | | **Interfaces per machine** | one · several | yes | `at:` takes a list | -| **Reachability policy** | symmetric · asymmetric | yes | `inbound:` — a routable machine that refuses everything | +| **Reachability policy, host** | symmetric · asymmetric | yes | `inbound:` — a routable machine that refuses everything | +| **Reachability policy, network** | open · segmented · asymmetric between segments | yes | `policy:` — inter-segment rules, as a segmented router enforces | +| **Segments per gateway** | one · several behind one router | yes | identical gateway declarations mean one gateway machine | | **Gateway state** | permanent · expiring mappings | yes | `mapping_ttl:` | | **Path MTU** | standard · reduced | yes | `segments[].mtu` | | **Overlapping ranges** | distinct · two sites both on `192.168.1.0/24` | yes | segments may carry the same range | @@ -357,11 +439,17 @@ revisited — and it would be revisited by a real failure, which is the right tr ### The honest summary -The model now covers **where a machine sits** and **what the path between machines is like**, -across both address families, which together are what the mesh's reachability logic turns on. +The model now covers **where a machine sits**, **what the path between machines is like**, and +**what is permitted between them** — across both address families. Together those are what the +mesh's reachability logic turns on. -What it does not model is *change over time* beyond moving a machine, and *degradation* short of -failure. Both are absences chosen rather than overlooked. +It contains almost no equipment, by the test above: a device that does not change what an IP +packet can do has no fault for a scenario to catch. What it does contain is every device that +does — which turns out to be the router, and a modem only when the modem is also a router. + +What it still does not model is *change over time* beyond moving a machine, *degradation* short +of failure, and *publishing through two gateways at once*. The first two are chosen; the third +is the one real absence, and it is exactly the double-NAT case. ## Open -- 2.54.0 From a87308814060bf7da7735f74f32b2e672e124f92 Mon Sep 17 00:00:00 2001 From: jochen Date: Sun, 23 Aug 2026 23:43:09 +0200 Subject: [PATCH 5/9] A worked example: the whole model applied to an ordinary mesh MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The shape research 004 identified — one machine with a routable address, one publicly named but behind a household connection, one stationary on that network, one that roams — written out with every field the model has, in role names and documentation addresses. It shows the ISP modem doing nothing, because in bridge mode it is a media converter: it changes the physical medium and leaves the packets alone, so it creates no IP-level fact and appears nowhere. In router mode it would be a second gateway and publishing would need a rule on both, which is the one case the model still cannot express. It shows two segments sharing one gateway declaration, which means one gateway machine, and a policy rule between them that is asymmetric because useful ones almost always are. And it shows what is deliberately absent. Research 004 recorded overlay addresses, hub election and names for exactly this topology, and none of them appear: a scenario must not state what the mesh is responsible for. Given the declaration, whether a hub is elected, whether the NATed machine's endpoint is learned, and whether the roaming machine re-forms after moving are all observed rather than arranged. The absence is the point. The run at the end moves one identity through four positions — home, foreign network, asleep, home again — against a foreign gateway whose mapping expires in 30 seconds, which is why a number is there rather than a boolean. --- 03-DESIGN/01-to-be/02-scenario-declaration.md | 109 ++++++++++++++++++ 1 file changed, 109 insertions(+) diff --git a/03-DESIGN/01-to-be/02-scenario-declaration.md b/03-DESIGN/01-to-be/02-scenario-declaration.md index 62e59b2..5789e4f 100644 --- a/03-DESIGN/01-to-be/02-scenario-declaration.md +++ b/03-DESIGN/01-to-be/02-scenario-declaration.md @@ -388,6 +388,115 @@ This is **not** the same as `machines[].inbound`, and conflating them loses a re A mesh node behind a `policy` deny cannot be reached even by a peer that knows exactly where it is — and that is a real topology, not a contrived one. +## Worked example — a whole mesh of the ordinary kind + +The shape research 004 identified: one machine with a routable address, one publicly named but +behind a household connection, one stationary machine on that network, one that roams. Written +out completely, with every field the model has. + +```yaml +scenario: the-ordinary-shape + +segments: + internet: + kind: public + cidr: [203.0.113.0/24, 2001:db8::/32] + mtu: 1500 + + home: # the household network + kind: private + cidr: [192.168.1.0/24, 2001:db8:1::/64] + mtu: 1492 # PPPoE on the uplink; 1500 if the line is not PPPoE + gateway: + to: internet + address: 203.0.113.50 # the address the ISP hands the household + nat: [v4] # v4 translated, v6 routed — the modern default + forwardable: true # the household router is ours to configure + mapping_ttl: 120s + + devices: # optional: a segmented network on the same router + kind: private + cidr: [192.168.30.0/24] + gateway: + to: internet + address: 203.0.113.50 # identical → the SAME gateway machine + nat: [v4] + forwardable: true + mapping_ttl: 120s + + elsewhere: # wherever the roaming machine happens to be + kind: private + cidr: [198.51.100.0/24] + mtu: 1400 + gateway: + to: internet + address: 203.0.113.80 + nat: [v4] + forwardable: false # someone else's network, or carrier-grade NAT + mapping_ttl: 30s + +policy: + - { from: devices, to: home, allow: false } + - { from: home, to: devices, allow: true } + +machines: + anchor: # routable, nothing in front of it + at: { segment: internet, address: [203.0.113.10, 2001:db8::10] } + inbound: allow + + home-server: # publicly named, behind the household connection + at: { segment: home, address: [192.168.1.135, 2001:db8:1::135] } + published: + - { port: 443, on: home } # v4 reaches it only through the forward + inbound: allow # and v6 reaches it directly, so this matters + + workstation: # on the household network, not published + at: { segment: home, address: [192.168.1.250, 2001:db8:1::250] } + inbound: deny + + laptop: # starts at home; moves during the run + at: { segment: home, address: [192.168.1.98, 2001:db8:1::98] } + inbound: deny + +place: + all: [host] + anchor: [substrate] + +snapshot: raised +``` + +### What the ISP modem is doing here + +**Nothing, if it is in bridge mode** — which is the ordinary arrangement when the household has +its own router. A bridged modem is a media converter: it changes the physical medium and leaves +the packets alone, so it creates no IP-level fact and appears nowhere above. + +Were it in router mode it would be a second gateway, `home` would sit behind it rather than +behind `internet` directly, and publishing would need a rule on both — the case the model +cannot yet express. + +### What is deliberately absent + +Nothing here mentions overlay addresses, which node is the hub, who peers with whom, any name, +or any certificate. Research 004 recorded all of those for this topology, and **a scenario must +not state them** ([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)): they +are what the mesh does, and a scenario that supplied them would be certifying its own work. + +The absence is the point. Given the declaration above, whether a hub is elected, whether the +NATed machine's endpoint is learned, whether the roaming machine re-forms after moving — all of +it is observed. + +### Running it + +``` +move laptop → { segment: elsewhere, address: [198.51.100.23] } # it leaves the house +move laptop → detached # it sleeps +move laptop → { segment: home, address: [192.168.1.98, 2001:db8:1::98] } +``` + +One identity, four positions, one run. The `mapping_ttl: 30s` on `elsewhere` means a connection +held without refreshing dies while it is out — which is the point of putting a number there. + ## Is this general? — the axes a setup can vary along The question that matters is not *does this cover our mesh*, but **can it express any mesh**. -- 2.54.0 From e88a6df9247343f2ba7b7161e944812e2f6bba0e Mon Sep 17 00:00:00 2001 From: jochen Date: Sun, 23 Aug 2026 23:49:38 +0200 Subject: [PATCH 6/9] Public networks are unrelated, and routed rather than bridged MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Caught in review: every public address sat in one /24, which made the three of them look like one network. They are not. The internet is a very large number of unrelated networks routing to each other, and a machine in one is many hops from a machine in another with no shared broadcast domain between them. Putting them in one prefix would have quietly made four false things true in the lab: machines resolving each other by ARP and talking directly, TTL never decrementing, broadcast and multicast crossing between them, and any two being adjacent. The third is not hypothetical. The as-is layer records that mesh names are deliberately not multicast names, after a delay and a one-node-only failure mode. A lab where the internet is one broadcast domain would let a node discover a peer by multicast that it could never discover in production, and report success — the exact false green this effort exists to prevent. So a scenario has one public segment per public NETWORK, each with its own unrelated prefix, wired together through a router and never onto a shared bridge. That is a property of how the lab wires them rather than a field anyone sets, because no correct scenario has two public networks adjacent. Addresses now spread across all three RFC 5737 ranges plus RFC 3849 /48s, chosen to look nothing like each other, and a foreign private network uses someone else's RFC 1918 range rather than a documentation one. All three examples in the document rewritten, since two of them still showed a single flat internet segment and contradicted the new rule. --- 03-DESIGN/01-to-be/02-scenario-declaration.md | 175 ++++++++++++------ 1 file changed, 120 insertions(+), 55 deletions(-) diff --git a/03-DESIGN/01-to-be/02-scenario-declaration.md b/03-DESIGN/01-to-be/02-scenario-declaration.md index 5789e4f..695268d 100644 --- a/03-DESIGN/01-to-be/02-scenario-declaration.md +++ b/03-DESIGN/01-to-be/02-scenario-declaration.md @@ -17,6 +17,54 @@ everything in the lab hangs off, so it is worth getting small. It states what a hosting provider and a home router would provide, and nothing the mesh is responsible for ([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)). +## Public networks are unrelated, and routed rather than bridged + +The internet is not a network. It is a very large number of unrelated networks that route to +each other, and a machine in one is **many hops** from a machine in another with no shared +broadcast domain between them. + +So a scenario does not have *an* internet segment. It has **one public segment per public +network**, each with its own unrelated prefix, and the lab wires them together **through a +router, never onto a shared bridge**. + +That distinction is load-bearing, and putting several public addresses in one prefix would +quietly make four things true that are false in reality: + +| If public addresses share a segment | Reality | +|---|---| +| machines resolve each other by ARP and talk directly | they are routed, many hops apart | +| TTL never decrements | every hop decrements it | +| broadcast and multicast reach across | neither crosses a router | +| any two are adjacent | adjacency is the exception, not the rule | + +The third is not hypothetical here. The mesh has already been bitten by multicast name +resolution — [`00-as-is/01-mesh-and-transport.md`](../00-as-is/01-mesh-and-transport.md) +records that mesh names are deliberately not multicast names, after a delay and a +one-node-only failure mode. A lab where "the internet" is one broadcast domain would let a node +discover a peer by multicast that it could never discover in production, and report success. + +**The rule: `kind: public` segments are routed to one another and never bridged.** It is a +property of how the lab wires them, not a field anyone sets, because there is no correct +scenario in which two public networks are adjacent. + +### Which addresses to use + +RFC 5737 reserves three ranges and RFC 3849 reserves one v6 prefix. Each public network takes +its own, and they are chosen to look nothing like each other — because in reality they would +not: + +| Public network | v4 | v6 | +|---|---|---| +| a hosting provider | `192.0.2.0/24` | `2001:db8:a::/48` | +| a household ISP | `198.51.100.0/24` | `2001:db8:b::/48` | +| a mobile or foreign network | `203.0.113.0/24` | `2001:db8:c::/48` | + +Three is enough for the topologies that matter, and a fourth public network can subnet one of +them — an ISP handing out `198.51.100.0/25` and `198.51.100.128/25` to two customers is exactly +what really happens. + +Private segments still use RFC 1918 and stay byte-identical to production. + ## What NAT does, and why the design turns on it A household or office has **one** address the outside world can see, and **many** machines @@ -103,53 +151,57 @@ The v6 documentation prefix is `2001:db8::/32` (RFC 3849) — the exact counterp scenario: roaming-and-published segments: - internet: - kind: public # RFC 5737 for v4, RFC 3849 for v6 - cidr: [203.0.113.0/24, 2001:db8::/32] + hosting: # one public network + kind: public + cidr: [192.0.2.0/24, 2001:db8:a::/48] + + isp-home: # another, unrelated — routed to it, not bridged + kind: public + cidr: [198.51.100.0/24, 2001:db8:b::/48] home: kind: private - cidr: [192.168.1.0/24, 2001:db8:1::/64] + cidr: [192.168.1.0/24, 2001:db8:b:1::/64] mtu: 1500 gateway: - to: internet - address: 203.0.113.50 # what the world sees this network as, on v4 + to: isp-home + address: [198.51.100.7, 2001:db8:b::7] # what the world sees this network as nat: [v4] # v4 is translated; v6 is routed, not translated forwardable: true mapping_ttl: 120s # an unused inbound mapping is forgotten after this - elsewhere: # a network we do not control + cafe: # a network we do not control kind: private - cidr: [198.51.100.0/24] # v4 only — no v6 offered here at all + cidr: [10.50.0.0/16] # v4 only — no v6 offered here at all mtu: 1400 # a tunnelled path, smaller than standard gateway: - to: internet - address: 203.0.113.80 + to: hosting # it reaches the world via a different public network + address: [192.0.2.200] nat: [v4] forwardable: false # café wifi, or carrier-grade NAT mapping_ttl: 30s # aggressive, as carrier NAT tends to be machines: anchor: - at: { segment: internet, address: [203.0.113.10, 2001:db8::10] } + at: { segment: hosting, address: [192.0.2.10, 2001:db8:a::10] } home-server: - at: { segment: home, address: [192.168.1.135, 2001:db8:1::135] } + at: { segment: home, address: [192.168.1.135, 2001:db8:b:1::135] } published: - - { port: 443, on: home } # v4 only: 203.0.113.50:443 → 192.168.1.135:443 + - { port: 443, on: home } # v4 only: 198.51.100.7:443 → 192.168.1.135:443 inbound: allow # v6 is routable here, so this decides whether it is reachable workstation: - at: { segment: home, address: [192.168.1.250, 2001:db8:1::250] } + at: { segment: home, address: [192.168.1.250, 2001:db8:b:1::250] } inbound: deny # a host firewall: dials out, accepts nothing laptop: - at: { segment: home, address: [192.168.1.98, 2001:db8:1::98] } + at: { segment: home, address: [192.168.1.98, 2001:db8:b:1::98] } border: # a machine on two segments at once at: - - { segment: home, address: [192.168.1.2] } - - { segment: internet, address: [203.0.113.60] } + - { segment: home, address: [192.168.1.2] } + - { segment: isp-home, address: [198.51.100.60] } place: all: [host] @@ -160,13 +212,14 @@ snapshot: raised ## What each part means, precisely -**`segments`** — a broadcast domain with an address range, and a `kind:`. +**`segments`** — a broadcast domain with an address range, and a `kind:`. A segment is a single +broadcast domain, which is precisely why several public networks cannot be one segment. -`kind: public` marks the segment that stands in for the internet. `kind: private` is everything -else. This is stated rather than inferred, and the earlier version inferred it — *a segment with +`kind: public` marks a segment that stands in for a public network — and there is normally more +than one, unrelated to each other. `kind: private` is everything else. This is stated rather than inferred, and the earlier version inferred it — *a segment with no gateway is the internet* — which made an **isolated network inexpressible**: a LAN with no -route out is a private segment with no gateway, and would have been read as the internet and -forced to use documentation addresses. A mesh spanning a site with no internet access is a real +route out is a private segment with no gateway, and would have been read as public and forced +to use documentation addresses. A mesh spanning a site with no internet access is a real topology, and the model has to be able to say it. **`gateway:`** — how a segment reaches its parent, and this is where the previous version was @@ -217,7 +270,7 @@ rather than a port list. `{ port: 443, on: home }` means the `home` gateway forw `203.0.113.50:443` to this machine's `443`. The resulting public endpoint is derivable, which is the point: a scenario never writes an endpoint down, and the mesh has to discover it. -A machine may be published on **any gateway between it and the internet** — which is how *"our +A machine may be published on **any gateway between it and a public network** — which is how *"our LAN also has a public IP"* is expressed, and why `on:` names the gateway rather than being implied. It cannot be published at all on a gateway the scenario models as foreign; attempting it is a declaration error, because that is precisely the constraint being reproduced. @@ -252,13 +305,8 @@ segment meant to be routable makes a would-be hub test as unreachable, and **the never forms** — no error, no failed step, just a mesh that does not exist. Research 004 calls this the single most important fact in its analysis. -RFC 5737 reserves three ranges, which is exactly enough for the topology above: - -| Range | Used for | -|---|---| -| `203.0.113.0/24` | the internet segment itself — directly attached machines, and gateway addresses | -| `198.51.100.0/24` | a foreign network, so a roaming machine's apparent address is plainly not ours | -| `192.0.2.0/24` | spare — a second foreign network, or a second site | +Which range goes where is covered above, under *public networks are unrelated*: one range per +public network, chosen to look nothing like each other. Private segments use RFC 1918 and can be **byte-identical to production**, because those addresses mean the same thing everywhere. Only the public side is substituted, and only because @@ -352,11 +400,11 @@ segments: trusted: kind: private cidr: [192.168.1.0/24] - gateway: { to: internet, address: 203.0.113.50, nat: [v4], forwardable: true } + gateway: { to: isp-home, address: [198.51.100.7], nat: [v4], forwardable: true } devices: kind: private cidr: [192.168.30.0/24] - gateway: { to: internet, address: 203.0.113.50, nat: [v4], forwardable: true } + gateway: { to: isp-home, address: [198.51.100.7], nat: [v4], forwardable: true } ``` Identical gateway declarations mean **one gateway machine**, not two. The lab materialises a @@ -398,18 +446,29 @@ out completely, with every field the model has. scenario: the-ordinary-shape segments: - internet: + # ---- three unrelated public networks. Routed to each other, never bridged. ---- + + hosting: # where the always-on machine lives kind: public - cidr: [203.0.113.0/24, 2001:db8::/32] - mtu: 1500 + cidr: [192.0.2.0/24, 2001:db8:a::/48] + + isp-home: # the household's uplink + kind: public + cidr: [198.51.100.0/24, 2001:db8:b::/48] + + isp-mobile: # wherever the roaming machine happens to be + kind: public + cidr: [203.0.113.0/24, 2001:db8:c::/48] + + # ---- private networks behind them ---- home: # the household network kind: private - cidr: [192.168.1.0/24, 2001:db8:1::/64] + cidr: [192.168.1.0/24, 2001:db8:b:1::/64] mtu: 1492 # PPPoE on the uplink; 1500 if the line is not PPPoE gateway: - to: internet - address: 203.0.113.50 # the address the ISP hands the household + to: isp-home + address: [198.51.100.7, 2001:db8:b::7] nat: [v4] # v4 translated, v6 routed — the modern default forwardable: true # the household router is ours to configure mapping_ttl: 120s @@ -418,21 +477,21 @@ segments: kind: private cidr: [192.168.30.0/24] gateway: - to: internet - address: 203.0.113.50 # identical → the SAME gateway machine + to: isp-home + address: [198.51.100.7, 2001:db8:b::7] # identical → the SAME gateway machine nat: [v4] forwardable: true mapping_ttl: 120s - elsewhere: # wherever the roaming machine happens to be + cafe: # a network we do not control kind: private - cidr: [198.51.100.0/24] + cidr: [10.50.0.0/16] # RFC 1918 — someone else's private range mtu: 1400 gateway: - to: internet - address: 203.0.113.80 + to: isp-mobile + address: [203.0.113.129] nat: [v4] - forwardable: false # someone else's network, or carrier-grade NAT + forwardable: false # carrier-grade NAT, or simply not ours mapping_ttl: 30s policy: @@ -441,21 +500,21 @@ policy: machines: anchor: # routable, nothing in front of it - at: { segment: internet, address: [203.0.113.10, 2001:db8::10] } + at: { segment: hosting, address: [192.0.2.10, 2001:db8:a::10] } inbound: allow home-server: # publicly named, behind the household connection - at: { segment: home, address: [192.168.1.135, 2001:db8:1::135] } + at: { segment: home, address: [192.168.1.135, 2001:db8:b:1::135] } published: - { port: 443, on: home } # v4 reaches it only through the forward inbound: allow # and v6 reaches it directly, so this matters workstation: # on the household network, not published - at: { segment: home, address: [192.168.1.250, 2001:db8:1::250] } + at: { segment: home, address: [192.168.1.250, 2001:db8:b:1::250] } inbound: deny laptop: # starts at home; moves during the run - at: { segment: home, address: [192.168.1.98, 2001:db8:1::98] } + at: { segment: home, address: [192.168.1.98, 2001:db8:b:1::98] } inbound: deny place: @@ -489,13 +548,18 @@ it is observed. ### Running it ``` -move laptop → { segment: elsewhere, address: [198.51.100.23] } # it leaves the house -move laptop → detached # it sleeps -move laptop → { segment: home, address: [192.168.1.98, 2001:db8:1::98] } +move laptop → { segment: cafe, address: [10.50.3.23] } # it leaves the house +move laptop → detached # it sleeps +move laptop → { segment: home, address: [192.168.1.98, 2001:db8:b:1::98] } ``` -One identity, four positions, one run. The `mapping_ttl: 30s` on `elsewhere` means a connection -held without refreshing dies while it is out — which is the point of putting a number there. +One identity, four positions, one run. The `mapping_ttl: 30s` on `cafe` means a connection held +without refreshing dies while it is out — which is the point of putting a number there. + +Note that the laptop's three positions are on **three different public networks**: at home it +appears as `198.51.100.7`, at the café as `203.0.113.129`, and the machine it is trying to reach +is on a fourth. None of them are adjacent, all of them are routed. That is the situation the +mesh actually faces. ## Is this general? — the axes a setup can vary along @@ -515,7 +579,8 @@ how the mesh behaves**. Bandwidth does not change correctness; MTU does. | **Gateway state** | permanent · expiring mappings | yes | `mapping_ttl:` | | **Path MTU** | standard · reduced | yes | `segments[].mtu` | | **Overlapping ranges** | distinct · two sites both on `192.168.1.0/24` | yes | segments may carry the same range | -| **Segment count** | one · many · isolated island | yes | `kind:` distinguishes an island from the internet | +| **Segment count** | one · many · isolated island | yes | `kind:` distinguishes an island from a public network | +| **Adjacency** | same segment · routed · unrelated public networks | yes | public segments are routed, never bridged — so no two are adjacent unless declared so | | **Gateway depth** | direct · one gateway · nested | **partly** | `to:` chains, so nesting exists; `published:` names one gateway, so forwarding through two does not | | **Address stability** | static · dynamic · changes mid-run | **partly** | a machine can be *moved*; an address changing under it in place cannot be stated | | **Path quality** | latency · loss · bandwidth | **no**, deliberately | changes performance, not correctness — modelling it makes a network simulator, not a fixture | -- 2.54.0 From a253afe020b9f9265f99db917fe7928b981e4f33 Mon Sep 17 00:00:00 2001 From: jochen Date: Sun, 23 Aug 2026 23:53:00 +0200 Subject: [PATCH 7/9] Scenario lifecycle, and how two scenarios coexist MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ADR 0032: a scenario is a closed address space. Every segment materialises as its own isolated link belonging to one instance, so two scenarios raised from the same declaration hold the same addresses and never meet. The declaration keeps its literal addresses and they mean what they say — allocating from a pool would have made them a fiction, so a scenario reproducing a specific topology would stop reproducing it. The constraint that follows shapes everything: the lab never reaches into a scenario over IP. It talks to machines through the virtualisation layer's own channel. If it reached them by address, the workstation would need a route into each scenario, and two carrying the same prefix would give it two routes to one destination — failing not with an error but by one scenario's traffic arriving in another. That also makes reachability an honest question. Can this machine reach that one is asked from INSIDE, by executing on the first, rather than probed from a workstation that is not on the network and whose opinion would be a different question with a misleadingly similar answer. The lifecycle itself: six verbs, of which raise and destroy are enough to be useful and the rest are what make repetition cheap. Raising is convergent rather than incremental, because a lab behaving differently from the thing it tests teaches the wrong habit. A failed raise leaves the wreckage standing. Tearing down on failure destroys the only evidence, which is backwards — a scenario that failed to raise is more interesting than one that succeeded. Snapshots are whole-scenario. Per-machine would be cheaper and wrong: the mesh keeps state spanning nodes, so restoring one machine while its peers move on produces a mesh that has never existed, and faults found there would be artefacts of the lab. Closes the declaration's open question about running several scenarios at once. --- ...a-scenario-is-an-isolated-address-space.md | 71 ++++++++++ 03-DESIGN/01-to-be/02-scenario-declaration.md | 4 - 03-DESIGN/01-to-be/03-scenario-lifecycle.md | 134 ++++++++++++++++++ 03-DESIGN/01-to-be/README.md | 1 + 4 files changed, 206 insertions(+), 4 deletions(-) create mode 100644 02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md create mode 100644 03-DESIGN/01-to-be/03-scenario-lifecycle.md diff --git a/02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md b/02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md new file mode 100644 index 0000000..09dac33 --- /dev/null +++ b/02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md @@ -0,0 +1,71 @@ +--- +status: accepted +date: 2026-08-23 +deciders: jochen +reconstructed: false +--- + +# 32. A scenario is an isolated address space, and the lab never reaches into it over IP + +## Context + +A scenario declares literal addresses — +[the declaration](../03-DESIGN/01-to-be/02-scenario-declaration.md) is full of them, and it has +to be, because reproducing *published but behind NAT* means saying which address the world sees. + +That raises a question the declaration left open: **two scenarios at once.** Several agents +working means several scenarios, and the lab design already calls that a requirement. But two +scenarios built from the same declaration want the same addresses, and there are only three +documentation ranges in existence. + +## Considered options + +1. **Allocate addresses from a pool at raise time**, rewriting the declaration's literals. + Rejected. It makes the addresses in a declaration a fiction, so a scenario reproducing a + specific topology no longer reproduces it; it breaks the RFC-range validation, since + allocated addresses would have to come from somewhere real; and the numbers a person reads + in the file stop being the numbers they will see in a capture. +2. **One scenario at a time.** Rejected — it is the requirement, not an inconvenience. A gate + an agent has to queue for is a gate that gets bypassed. +3. **Give each scenario its own network stack, so the addresses do not collide.** Chosen. + +## Decision + +**A scenario is a closed address space.** Every segment materialises as its own isolated link, +belonging to one scenario instance. Two scenarios raised from the same declaration hold the same +addresses and never meet, because nothing joins their links. + +The declaration therefore keeps its literal addresses, and they mean exactly what they say. + +**The consequence that constrains everything else: the lab never reaches into a scenario over +IP.** It talks to a machine through the virtualisation layer's own channel — the same way one +executes a command in a container without the container being routable. + +That is not a preference. If the lab reached machines by address, the workstation running it +would need a route into each scenario, and two scenarios carrying the same prefix would give it +two routes to the same destination. Concurrency would be impossible, and it would fail in the +worst available way: not with an error, but by one scenario's traffic arriving in another. + +## Consequences + +- Scenarios are concurrent by construction, with no allocation, no bookkeeping and no limit + beyond the machine's capacity. +- The three documentation ranges stop being a scarce resource. Every scenario may use all of + them, because no two scenarios share a link. +- **The lab cannot use IP to check anything**, which is more of a constraint than it first + appears: *"can this machine reach that one"* has to be asked **from inside the scenario**, by + executing on a machine, rather than probed from outside. That is the honest way to ask it + anyway — reachability from the workstation is not the question. +- A scenario is a unit that can be paused, snapshotted and destroyed whole, because nothing + outside holds a reference into it. +- The lab needs a scenario **instance** identity distinct from the scenario name in the + declaration: the declaration is a kind, and several instances of one kind may exist. +- **The workstation is not on the scenario's network, so it is not a node in it.** Anything a + developer wants to reach — a web interface, a database — needs an explicit, deliberate + forward out of the scenario, which is a feature rather than a gap: nothing leaks by default. + +## References + +- [ADR 0031](0031-the-lab-provides-the-underlay.md) — the declaration whose literal addresses + this preserves. +- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the lifecycle jobs this shapes. diff --git a/03-DESIGN/01-to-be/02-scenario-declaration.md b/03-DESIGN/01-to-be/02-scenario-declaration.md index 695268d..370be64 100644 --- a/03-DESIGN/01-to-be/02-scenario-declaration.md +++ b/03-DESIGN/01-to-be/02-scenario-declaration.md @@ -634,10 +634,6 @@ is the one real absence, and it is exactly the double-NAT case. - **Where `place:` gets its artifacts from.** Before the mesh is self-hosting these come from outside; afterwards from the mesh itself. The declaration should not have to care, which suggests a named source rather than a path. -- **Multiple scenarios at once.** Each needs its own segments and addresses, and the shape above - writes addresses absolutely. Whether a scenario carries literal addresses or a template the - lab allocates from decides whether two can run side by side — and there are only three - documentation ranges to go round. - **Nested forwarding** — `published:` names one gateway, so a machine behind two cannot be published through both. - **An address changing in place**, as a DHCP lease expiring under a machine that has not moved. diff --git a/03-DESIGN/01-to-be/03-scenario-lifecycle.md b/03-DESIGN/01-to-be/03-scenario-lifecycle.md new file mode 100644 index 0000000..3ca2623 --- /dev/null +++ b/03-DESIGN/01-to-be/03-scenario-lifecycle.md @@ -0,0 +1,134 @@ +--- +layer: to-be +status: designed +code: [mesh-lab] +updated: 2026-08-23 +decisions: + - 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md + - 02-DECISIONS/0031-the-lab-provides-the-underlay.md + - 02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md +--- + +# Scenario lifecycle + +The first thing the lab must do, and the only thing it must do before anything else can be +written: **materialise a mesh, return it to a known state, and destroy it** +([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)). + +A [declaration](02-scenario-declaration.md) describes a scenario. This describes what happens +to one. + +## The verbs + +| Verb | Does | +|---|---| +| `raise` | materialise a declaration into a running scenario instance | +| `snapshot` | name the current state of the whole scenario | +| `restore` | return the whole scenario to a named state | +| `move` | change a machine's position while the scenario runs | +| `exec` | run something on a machine, and get its output | +| `destroy` | tear the instance down | + +Six verbs, and `raise` plus `destroy` are enough to be useful. The rest are what make repetition +cheap, and cheap repetition is what turns the bootstrap path into an inner development loop +rather than a ceremony. + +## Raising, in order + +The order is not arbitrary — each step needs the one before it to exist: + +1. **Segments.** Isolated links, one per declared segment, belonging to this instance and + joined to nothing outside it + ([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)). +2. **Gateways.** Derived, never declared as machines: a gateway is materialised for each + distinct `gateway:` declaration, sitting on both its segment and its parent, carrying the + translation, forwarding and mapping-expiry the declaration asked for. +3. **Machines.** Each on its segments, holding its addresses. +4. **Policy.** Rules between segments, applied on the gateways that route between them. +5. **Placement.** Artifacts onto machines. +6. **Snapshot**, if the declaration named one. + +Raising is **convergent, not incremental**: raising an instance that already exists brings it to +the declared state rather than failing or duplicating. That is the same model the mesh itself +uses, and a lab that behaved differently from the thing it tests would be teaching the wrong +habit. + +## A failed raise leaves the wreckage + +A step that fails stops the raise +([ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md)) — and **does not tear +down**. + +Tearing down on failure destroys the only evidence of what went wrong, which is precisely +backwards: a scenario that failed to raise is more interesting than one that succeeded. The +instance stays, marked failed, with the step that failed named. + +The caller decides what happens next, and the two callers want different things +— the coordinator captures and destroys, a person opens a shell. That is the same +one-runner-two-callers split the lab design already makes, applied to failure. + +## Snapshots are whole-scenario + +A snapshot captures **every machine and the state of the network between them**, as one thing. +Restoring returns all of it. + +Per-machine snapshots would be cheaper and are wrong. The mesh keeps state that spans nodes — +what is assigned where, which grants exist, what has been delivered — so restoring one machine +to an earlier moment while its peers move on produces a mesh that has never existed and could +not. The faults found there would be artefacts of the lab. + +This is what makes *fresh* and *upgrade* both cheap and both default: one snapshot of a mesh +that has never seen a change, another of a mesh running the previous version, and a restore +between runs. + +## Moving a machine + +`move` changes a machine's position while the scenario runs: to another segment, with different +addresses, or to `detached`. + +It is the roaming case, and it is a **lifecycle** operation rather than a declaration because +the interesting part is the transition, not the destination. A mesh that forms correctly with a +node at home and correctly with it away may still fail to notice it moved. + +Moving does not invalidate a snapshot. A snapshot is a state to return to; a move is a change +made after it, and returning undoes it like any other change. + +## Reaching in + +Everything the lab does to a machine goes through the virtualisation layer, never over IP +([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)). `exec` runs a +command on a machine and returns its output. + +This has one consequence worth stating plainly: **a reachability question is asked from inside**. +*Can this machine reach that one* is `exec` on the first, testing the second — not a probe from +the workstation. The workstation is not on the scenario's network and its opinion of +reachability would be a different question with a misleadingly similar answer. + +Anything a person wants to open in a browser needs a deliberate forward out of the instance. +Nothing leaks by default. + +## The two callers + +The lab design already establishes that the runner serves the coordinator and a person, and +that anything only one of them can do will drift. Applied here: + +| | the coordinator | someone working on the mesh | +|---|---|---| +| wants | structured results, clean teardown, no prompts | readable output, the scenario **left standing**, one thing re-run without repeating the rest | +| on failure | capture, then destroy | leave it, open a shell | + +Both use the same verbs. The difference is what happens after the verdict, which is a caller's +decision rather than a second implementation. + +## Open + +- **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the + inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop + is slow, and everything above is theory. +- **Instance naming.** A declaration is a kind and instances are many; how they are named + decides whether a person can find the one they left standing yesterday. +- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying + the instance must not destroy them. +- **Placement before the mesh is self-hosting.** `place:` needs artifacts from somewhere, and + before the mesh builds itself that somewhere is outside it — the open question from + [research 006](../../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md). diff --git a/03-DESIGN/01-to-be/README.md b/03-DESIGN/01-to-be/README.md index e46578d..70e82bf 100644 --- a/03-DESIGN/01-to-be/README.md +++ b/03-DESIGN/01-to-be/README.md @@ -12,6 +12,7 @@ document is written and this one's status becomes `implemented`. | [`00-work-breakdown.md`](00-work-breakdown.md) | How the decomposition gets built, in what order, and where a human must look | [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) | | [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md), [0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) | | [`02-scenario-declaration.md`](02-scenario-declaration.md) | What a scenario declares — the underlay, and what to place on it | [ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md) | +| [`03-scenario-lifecycle.md`](03-scenario-lifecycle.md) | What happens to a scenario — raise, snapshot, restore, move, destroy | [ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md) | ## Not yet written -- 2.54.0 From 98bcd5cc4941cc6475a6e4a9823442a53995b6ef Mon Sep 17 00:00:00 2001 From: jochen Date: Mon, 24 Aug 2026 00:07:22 +0200 Subject: [PATCH 8/9] =?UTF-8?q?Measure=20the=20lab's=20inner=20loop=20?= =?UTF-8?q?=E2=80=94=20it=20is=20too=20slow,=20for=20a=20fixable=20reason?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The lifecycle design closed on a question that was measurable rather than arguable, so it was measured. One virtual machine on a workstation with hardware virtualisation and NVMe. Raising: the launch call returns in 3.4s, the machine is actually usable after 14.3s. The gap is a design constraint — raise must wait for the second number, because reporting the first would be transport reported as effect, which is the mesh's own recurring failure. Snapshot: 9.9s and 1.6 GB for a 1.5 GB instance. A dir snapshot is a full copy; nothing is shared. Restore: 10.4s, usable again after 20.1s. The second snapshot exceeded two minutes and never completed. That is the more troubling number: snapshot cost here is not merely high, it is unpredictable, and a loop with a variable multi-minute step is one nobody trusts. Projected to a four-machine scenario, a reset-and-rerun cycle is about a minute and a half at best and unbounded at worst, before any of the mesh's own work begins. That is too slow for an inner loop, and ADR 0029's whole argument — that making the bootstrap path the inner loop turns the least-exercised code into the most-exercised — holds only while resetting is cheap. The cause is not virtual machines. Hardware virtualisation is present and machines boot in fourteen seconds. It is that the daemon offers exactly one storage driver, dir, which has no copy-on-write and therefore no cheap snapshot. The btrfs kernel module is available; btrfs-progs is simply not installed, which is the entire reason the driver is absent. The copy-on-write comparison was deliberately NOT run, because running it would mean installing a package by hand — which the rules forbid and which would have made the measurement unreproducible. So the honest statement is that the current configuration is too slow and the likely fix is known but unverified, rather than that btrfs fixes it. --- .../010-lab-inner-loop-cost/00-overview.md | 55 +++++++++++ .../010-lab-inner-loop-cost/measurements.md | 98 +++++++++++++++++++ 03-DESIGN/01-to-be/03-scenario-lifecycle.md | 12 ++- 3 files changed, 162 insertions(+), 3 deletions(-) create mode 100644 01-RESEARCH/010-lab-inner-loop-cost/00-overview.md create mode 100644 01-RESEARCH/010-lab-inner-loop-cost/measurements.md diff --git a/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md b/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md new file mode 100644 index 0000000..bd52bde --- /dev/null +++ b/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md @@ -0,0 +1,55 @@ +--- +status: active +initiated: 2026-08-24 +touches: + - 03-DESIGN/01-to-be/03-scenario-lifecycle.md + - 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md + - 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md +became: [] +--- + +# 010 — What the lab's inner loop actually costs + +## What is being investigated + +[The lifecycle design](../../03-DESIGN/01-to-be/03-scenario-lifecycle.md) closes on an open +question that is measurable rather than arguable: + +> **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the +> inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop +> is slow, and everything above is theory. + +Measured on a workstation, 2026-08-24. Numbers in [`measurements.md`](measurements.md). + +## Why it matters + +[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) makes the +bootstrap scenario the inner development loop for tiers 0 and 1 — the argument being that +raising a node from nothing stops being the least-exercised path and becomes the most-exercised +one. **That argument is only true if raising and resetting are cheap.** A loop that costs +minutes is a loop people avoid, and the least-exercised path stays least-exercised. + +## Status + +Measured, and the finding is a blocker rather than a data point. + +**A snapshot on the current host is a full copy of the machine's disk.** 1.6 GB and ten seconds +for one small virtual machine at best — and over two minutes when observed a second time. Cost +scales with the number of machines and the size of their disks, not with what changed. + +The cause is not virtual machines and not incus. It is that the host offers incus exactly one +storage driver, `dir`, which has no copy-on-write and therefore no cheap snapshot. The kernel +supports btrfs; the userspace tool that would let incus use it is simply not installed. + +So the question *"is the lab's inner loop fast enough"* currently answers itself the wrong way, +for a reason that is one declared package away from being fixed — and declaring packages is +something the mesh already does. + +## Open questions + +| Question | Why it matters | +|---|---| +| How much does a copy-on-write pool actually improve it? Expected to be near-instant snapshots and delta-sized storage, but **expected is not measured**. | The whole inner-loop argument rests on the answer. | +| Why was the second snapshot more than twelve times slower than the first? | If snapshot cost is unpredictable rather than merely high, that is worse — a loop with a variable multi-minute step is one nobody trusts. | +| Does a scenario snapshot need the machines stopped? | Stateless snapshots of a running virtual machine capture the disk but not memory. Whether a mesh restored that way is coherent is not established. | +| What is the cost at scenario scale — four machines rather than one? | Only single-machine numbers were taken. If the operation is serial, four machines is four times the wait. | diff --git a/01-RESEARCH/010-lab-inner-loop-cost/measurements.md b/01-RESEARCH/010-lab-inner-loop-cost/measurements.md new file mode 100644 index 0000000..9189c07 --- /dev/null +++ b/01-RESEARCH/010-lab-inner-loop-cost/measurements.md @@ -0,0 +1,98 @@ +--- +effort: 010-lab-inner-loop-cost +updated: 2026-08-24 +--- + +# Measurements + +Taken 2026-08-24 on a workstation with hardware virtualisation available, an NVMe-backed ext4 +root, and 300 GB free. One virtual machine, 1 GiB memory, 2 CPUs, from a cached distribution +image. + +## The environment, before anything ran + +| Fact | Value | Consequence | +|---|---|---| +| Hardware virtualisation | present | virtual machines run at native speed; the choice in [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md) is not paying an emulation penalty | +| Storage drivers the daemon offers | **`dir` only** | no copy-on-write, therefore no cheap snapshot | +| Host filesystems | ext4 throughout | nothing copy-on-write to put a pool on | +| btrfs kernel module | **available** | the kernel can do it | +| `btrfs-progs` | **not installed** | which is the entire reason the driver is absent | + +The last two rows are the finding. The daemon advertises only `dir` because the userspace tool +for anything better is missing — not because the host cannot do better. + +## Raising a machine + +| Step | Time | +|---|---| +| launch call returns | **3.4 s** | +| machine actually usable — a command executes on it | **14.3 s** | + +The gap matters for the lifecycle design: `raise` returning is not the same as the scenario +being ready, so the verb has to wait for the second number, not report the first. Reporting +the first would be the mesh's own recurring failure — transport reported as effect. + +## Snapshot and restore + +| Operation | Time | Disk | +|---|---|---| +| snapshot, first | **9.9 s** | **+1.6 GB** | +| snapshot, second | **> 120 s — did not complete** | — | +| restore call returns | **10.4 s** | — | +| machine usable again | **20.1 s** total | — | + +Instance on disk before snapshotting: 1.5 GB. Snapshot directory afterwards: 1.6 GB. **A `dir` +snapshot is a full copy** — the storage cost equals the instance, and nothing is shared. + +Implied copy throughput on the first snapshot is roughly 160 MB/s, which is far below what the +underlying NVMe can do and is consistent with a real, durable copy rather than a metadata +operation. + +**The second snapshot is the more troubling number.** It exceeded two minutes and was still +running when the observation was cut off; only the first snapshot exists. Whatever the cause — +page cache exhausted by the preceding restore, writeback contention — the practical +consequence is that snapshot cost here is **not merely high, it is unpredictable**. + +## What this projects to + +A four-machine scenario, taking the optimistic single-machine numbers and assuming the +operations are serial: + +| | one machine | four machines | +|---|---|---| +| raise, to usable | 14 s | ~57 s | +| snapshot | 10 s, 1.6 GB | ~40 s, 6.4 GB | +| restore, to usable | 20 s | ~80 s | + +A reset-and-rerun cycle is therefore **around a minute and a half at best**, and unbounded at +worst, before any of the mesh's own work begins. + +## The judgement + +**This is too slow for an inner loop**, and the reason is not the design. + +[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) argues that +making the bootstrap path the inner development loop turns the least-exercised code in the +system into the most-exercised. That argument holds only while resetting is cheap. At a minute +and a half a cycle, with occasional multi-minute stalls, the loop is one a person works around +— and the path stays under-exercised for exactly the reason it always was. + +Nothing about virtual machines causes this. Hardware virtualisation is present and the machines +boot in fourteen seconds. **The cost is entirely the storage driver**, and the driver is absent +because one userspace package is not installed on the host. + +The mesh already has the mechanism for that: a module declares a package, and a hook makes it a +working capability — which is precisely what was just done for the virtualisation daemon +itself, and what [`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md) +is about. + +## Not measured, and it matters + +The copy-on-write comparison **was not run**, because running it would mean installing a package +by hand, which the mesh's rules forbid and which would have made the measurement unreproducible +anyway. Copy-on-write snapshots are expected to be near-instant with storage proportional to +what changed. That expectation is well founded and is still an expectation. + +Until it is measured, the correct statement is: *the current configuration is too slow, and the +likely fix is known but unverified.* diff --git a/03-DESIGN/01-to-be/03-scenario-lifecycle.md b/03-DESIGN/01-to-be/03-scenario-lifecycle.md index 3ca2623..8a06df0 100644 --- a/03-DESIGN/01-to-be/03-scenario-lifecycle.md +++ b/03-DESIGN/01-to-be/03-scenario-lifecycle.md @@ -48,6 +48,10 @@ The order is not arbitrary — each step needs the one before it to exist: 5. **Placement.** Artifacts onto machines. 6. **Snapshot**, if the declaration named one. +`raise` must wait for the machines to be **usable**, not for the call to return. Measured, those +are 14 seconds apart for a single machine — and reporting the earlier number would be the mesh's +own recurring failure, transport reported as effect. + Raising is **convergent, not incremental**: raising an instance that already exists brings it to the declared state rather than failing or duplicating. That is the same model the mesh itself uses, and a lab that behaved differently from the thing it tests would be teaching the wrong @@ -122,9 +126,11 @@ decision rather than a second implementation. ## Open -- **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the - inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop - is slow, and everything above is theory. +- **What a snapshot costs — measured, and the answer is a blocker.** On the current host a + snapshot is a full copy: 1.6 GB and ten seconds for one small machine, over two minutes when + observed again, projecting to roughly a minute and a half for a four-machine reset-and-rerun + cycle. The cause is the storage driver, not virtual machines. See + [research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md). - **Instance naming.** A declaration is a kind and instances are many; how they are named decides whether a person can find the one they left standing yesterday. - **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying -- 2.54.0 From e88b448145585957b2ddbd000ceb3d908915c185 Mon Sep 17 00:00:00 2001 From: jochen Date: Mon, 24 Aug 2026 00:14:57 +0200 Subject: [PATCH 9/9] The fix is real: 76x, verified. And how the lab installs on a clean machine MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Snapshot 9.9s -> 0.13s. Restore 10.4s -> 0.80s. Three snapshots sharing 1.36 GB instead of costing 4.8 GB. The projected four-machine reset cycle falls from ~90s, unbounded at worst, to ~15s dominated by a boot that cannot be avoided. ADR 0029's inner-loop argument holds with copy-on-write and did not without it. The consistency matters as much as the speed: three consecutive snapshots took 0.13, 0.12 and 0.13 seconds, against a dir second snapshot that never finished. One honest counter-observation recorded: launching onto the fresh copy-on-write pool was slower, 20.2s against 14.3s, because the image had to be unpacked into a pool that had never seen it. Paid once per pool, and dwarfed by what snapshotting saves, but it went the other way. Doing the measurement produced the answer to how the lab installs on a clean machine, because both failure modes appeared while doing it. Installed is not available: the daemon was present with units disabled and no group. Issue 007. Available is not adequate, and this is worse: with the storage tooling absent everything worked and snapshots were seventy-six times slower. Nothing failed, nothing warned. That is a variant the mesh has not catalogued — its usual failure is reported success and did nothing; this is reported success and did it seventy-six times slower, which no error surface catches because nothing is wrong. So the lab verifies CAPABILITY, never installation, and refuses to run degraded rather than warning — a warning about a slow inner loop is read once and ignored forever. Prerequisites may arrive from a mesh module or from the lab's own bootstrap, and the second path is required rather than convenient: a lab installable only by a mesh cannot host the development of the mesh that installs it. The lab is the second thing installed by hand, after the node host, and for the same reason: something has to be first, and pretending otherwise produces a circularity papered over by a script nobody exercises. --- .../010-lab-inner-loop-cost/00-overview.md | 7 +- .../010-lab-inner-loop-cost/measurements.md | 57 +++++++-- 03-DESIGN/01-to-be/04-lab-installation.md | 116 ++++++++++++++++++ 03-DESIGN/01-to-be/README.md | 1 + 4 files changed, 173 insertions(+), 8 deletions(-) create mode 100644 03-DESIGN/01-to-be/04-lab-installation.md diff --git a/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md b/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md index bd52bde..db40f90 100644 --- a/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md +++ b/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md @@ -37,6 +37,10 @@ Measured, and the finding is a blocker rather than a data point. for one small virtual machine at best — and over two minutes when observed a second time. Cost scales with the number of machines and the size of their disks, not with what changed. +**With copy-on-write it is 0.13 seconds and costs the delta.** Verified, not assumed. The +projected four-machine reset cycle falls from roughly ninety seconds to roughly fifteen, of +which almost all is a boot that cannot be avoided. + The cause is not virtual machines and not incus. It is that the host offers incus exactly one storage driver, `dir`, which has no copy-on-write and therefore no cheap snapshot. The kernel supports btrfs; the userspace tool that would let incus use it is simply not installed. @@ -49,7 +53,8 @@ something the mesh already does. | Question | Why it matters | |---|---| -| How much does a copy-on-write pool actually improve it? Expected to be near-instant snapshots and delta-sized storage, but **expected is not measured**. | The whole inner-loop argument rests on the answer. | +| ~~How much does a copy-on-write pool actually improve it?~~ **Measured: snapshot 9.9 s → 0.13 s, restore 10.4 s → 0.80 s, three snapshots sharing 1.36 GB rather than costing 4.8 GB.** The projected four-machine cycle falls from ~90 s to ~15 s. | Answered. The inner-loop argument holds *with* copy-on-write and did not without it. | +| How does the lab install its own prerequisites on a clean machine? | The lab needs a virtualisation daemon, copy-on-write tooling and a pool before it can do anything — and it cannot depend on the mesh for them, since it is where the mesh is built. | | Why was the second snapshot more than twelve times slower than the first? | If snapshot cost is unpredictable rather than merely high, that is worse — a loop with a variable multi-minute step is one nobody trusts. | | Does a scenario snapshot need the machines stopped? | Stateless snapshots of a running virtual machine capture the disk but not memory. Whether a mesh restored that way is coherent is not established. | | What is the cost at scenario scale — four machines rather than one? | Only single-machine numbers were taken. If the operation is serial, four machines is four times the wait. | diff --git a/01-RESEARCH/010-lab-inner-loop-cost/measurements.md b/01-RESEARCH/010-lab-inner-loop-cost/measurements.md index 9189c07..a690054 100644 --- a/01-RESEARCH/010-lab-inner-loop-cost/measurements.md +++ b/01-RESEARCH/010-lab-inner-loop-cost/measurements.md @@ -87,12 +87,55 @@ working capability — which is precisely what was just done for the virtualisat itself, and what [`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md) is about. -## Not measured, and it matters +## The fix, measured -The copy-on-write comparison **was not run**, because running it would mean installing a package -by hand, which the mesh's rules forbid and which would have made the measurement unreproducible -anyway. Copy-on-write snapshots are expected to be near-instant with storage proportional to -what changed. That expectation is well founded and is still an expectation. +The comparison was subsequently run. One package — `btrfs-progs`, no dependencies — installed by +hand, the daemon restarted so it re-detected drivers, a copy-on-write pool created on a loop +file, and the identical image launched onto it. -Until it is measured, the correct statement is: *the current configuration is too slow, and the -likely fix is known but unverified.* +| Operation | `dir` | copy-on-write | | +|---|---|---|---| +| snapshot | 9.9 s, then **> 120 s** | **0.13 s** | ~76× faster, and *consistent* | +| snapshot again | — | 0.12 s | | +| snapshot a third time | — | 0.13 s | | +| restore call | 10.4 s | **0.80 s** | ~13× faster | +| restore, to usable | 20.1 s | **10.5 s** | the remainder is boot, which is irreducible | +| three snapshots, storage | ~4.8 GB | **1.36 GB total, shared** | cost is the delta, not the disk | + +**The fix is real, and larger than expected.** Snapshot goes from ten seconds to a tenth of a +second, and — more importantly — from *wildly variable* to *flat*. Three consecutive snapshots +took 0.13, 0.12 and 0.13 seconds. On `dir` the second snapshot never finished. + +Storage stops scaling with the machine and starts scaling with what changed: three snapshots of +a 1.5 GB instance occupied 1.36 GB in total, because they share. + +### What it projects to + +A four-machine reset-and-rerun cycle, the operation the inner loop repeats most: + +| | `dir` | copy-on-write | +|---|---|---| +| snapshot the scenario | ~40 s, 6.4 GB | **~0.5 s**, delta-sized | +| restore it | ~40 s + boot | **~3 s** + boot | +| **cycle** | **~90 s, unbounded at worst** | **~15 s, dominated by boot** | + +At fifteen seconds, dominated by a boot that cannot be avoided, the inner loop is viable and +[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)'s argument holds. +At ninety it did not. + +### One honest counter-observation + +Launching onto the fresh copy-on-write pool was **slower** — 20.2 s to usable against 14.3 s — +because the image had to be unpacked into a pool that had never seen it. That cost is paid once +per pool, not per scenario, and it is dwarfed by what snapshotting saves. But it is a real +number and it went the other way. + +### State this left behind + +Recorded because hand-made state is exactly what the mesh's rules exist to prevent, and it must +be declared properly rather than left as an artefact of a measurement: + +- `btrfs-progs` installed by hand. Its installation regenerated the boot initramfs, a side + effect worth knowing about. +- The daemon restarted once, to re-detect drivers. +- The test pool and instance were **removed**; the pool the lab actually needs does not exist. diff --git a/03-DESIGN/01-to-be/04-lab-installation.md b/03-DESIGN/01-to-be/04-lab-installation.md new file mode 100644 index 0000000..c1bd0cd --- /dev/null +++ b/03-DESIGN/01-to-be/04-lab-installation.md @@ -0,0 +1,116 @@ +--- +layer: to-be +status: designed +code: [mesh-lab] +updated: 2026-08-24 +decisions: + - 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md + - 02-DECISIONS/0008-a-failed-step-fails-the-job.md +--- + +# Installing the lab on a clean machine + +The lab has prerequisites — a virtualisation daemon, copy-on-write storage, a pool, an identity +permitted to talk to it — and it cannot get them from the mesh, because it is where the mesh is +built ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)). + +So the lab needs an install path of its own. This describes it, and the shape it has to take is +determined by two failures observed while measuring +([research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md)). + +## The two failures that shape this + +**One: installed is not available.** The virtualisation package was present and explicitly +installed. Both its units were disabled, the operator was in no group, and the client reported +the server unreachable. Nothing had failed — the declaration was satisfied exactly as written +([`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)). + +**Two, and worse: available is not adequate.** With the storage tooling absent, the daemon +offered one driver, everything worked, and snapshots took **seventy-six times longer** than they +needed to. Nothing failed. Nothing warned. A lab in that state runs correctly and is simply too +slow to use — and the inner loop it exists to provide quietly does not exist. + +The second is the more dangerous shape and it is a variant the mesh has not catalogued before. +Its usual failure is *reported success and did nothing*. This is **reported success and did it +seventy-six times slower**, which no error surface catches because nothing is wrong. + +## What follows: the lab verifies capability, never installation + +The install path may differ. **The verification does not.** + +Before the lab will raise anything, it asserts the outcomes it needs — not that packages are +present, but that the machine can actually do the work: + +| Assertion | Failing means | +|---|---| +| the daemon answers **as the invoking user**, not as root | a group membership that was granted but never took effect | +| a copy-on-write storage driver is offered | the userspace tooling is missing; snapshots will be full copies | +| **the pool the lab will use is on that driver** | a pool exists, and is the slow kind — the failure that has no symptom | +| hardware virtualisation is present | machines will be emulated and unusably slow | +| an image can be fetched or is cached | the first raise will fail late instead of early | + +Each check states **why it matters**, in the terms of what it costs. *"The pool uses the `dir` +driver"* means nothing to someone who does not already know it means seventy-six times slower +and unbounded at worst. + +**The lab refuses to run degraded.** It does not warn and continue: a warning about a slow inner +loop is read once and ignored forever, and the loop stays slow. This is +[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) applied where the failure is +performance rather than an error. + +## Two ways the prerequisites arrive + +**On a machine the mesh manages** — a module declares them, and a hook turns them into +capabilities: the units enabled, the group granted, the pool created on the right driver. This +already works; it is what was done for the virtualisation daemon itself. + +**On a machine the mesh does not manage** — the lab's own bootstrap does it. One command, on a +clean machine, that installs what is missing and configures it. + +The second path is not a convenience. It is **required**, because the lab must work before the +mesh does, and a lab that could only be installed by a mesh would be unable to host the +development of the mesh that installs it. + +Both paths end at the same verification. Whoever satisfied the prerequisites, the lab checks +them itself — because the lesson of the first failure is precisely that *something else said it +was done* is not evidence. + +## The lab is the second thing installed by hand + +Worth stating, because it looks like an exception and is not. + +The node host is the one thing installed by hand on a machine +([research 006](../../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)): everything else +arrives through it. The lab is the same shape on a workstation — installed once, by hand, and +then everything about the mesh is developed inside it. + +Two bootstraps, at two levels, for the same reason: **something has to be first, and pretending +otherwise produces a circularity that gets papered over with a script nobody exercises.** + +## What a clean install actually needs + +In order, on a machine with nothing: + +1. **A virtualisation daemon**, running, with its socket enabled. +2. **Copy-on-write storage tooling** — the kernel side is usually already present; it is the + userspace half that is missing and that decides whether the driver is offered at all. +3. **A pool on that driver.** A loop-backed file is sufficient and needs no partitioning, which + matters: requiring a dedicated filesystem would make the lab uninstallable on a machine + already in use. +4. **Group membership** for the operator — which does **not** apply to sessions that already + existed. Observed directly: a shell whose process tree predated the grant could not reach + the daemon while a fresh lookup showed the membership present. The bootstrap has to say so, + or the first thing a person meets is a permission error that looks like a broken install. +5. **Verification**, as above, before anything is raised. + +## Open + +- **Whether the lab's bootstrap may install packages at all**, given that the mesh's rules + forbid installing by hand. The resolution is probably that the lab's bootstrap *is* the + sanctioned mechanism on an unmanaged machine, in the way the mesh's own first-node script is — + but that is an argument to record, not to assume. +- **What "adequate" means numerically.** The checks above are qualitative. A snapshot-time + threshold would catch a copy-on-write pool that is slow for some other reason, and would be a + real assertion rather than a proxy. +- **Whether the lab should own its pool** rather than using an existing one. Owning it makes the + driver guaranteed; sharing it avoids duplicating storage on a machine that already has a pool. diff --git a/03-DESIGN/01-to-be/README.md b/03-DESIGN/01-to-be/README.md index 70e82bf..7f2fcc7 100644 --- a/03-DESIGN/01-to-be/README.md +++ b/03-DESIGN/01-to-be/README.md @@ -13,6 +13,7 @@ document is written and this one's status becomes `implemented`. | [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md), [0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) | | [`02-scenario-declaration.md`](02-scenario-declaration.md) | What a scenario declares — the underlay, and what to place on it | [ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md) | | [`03-scenario-lifecycle.md`](03-scenario-lifecycle.md) | What happens to a scenario — raise, snapshot, restore, move, destroy | [ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md) | +| [`04-lab-installation.md`](04-lab-installation.md) | Getting the lab onto a clean machine, and why it verifies capability rather than installation | [ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) | ## Not yet written -- 2.54.0