The scenario model, the lifecycle, and what the lab actually costs #6
Merged
jschoubben
merged 9 commits from 2026-08-23 22:40:55 +00:00
design/scenario-underlay-detail into main
9
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e88b448145 |
The fix is real: 76x, verified. And how the lab installs on a clean machine
Snapshot 9.9s -> 0.13s. Restore 10.4s -> 0.80s. Three snapshots sharing 1.36 GB instead of costing 4.8 GB. The projected four-machine reset cycle falls from ~90s, unbounded at worst, to ~15s dominated by a boot that cannot be avoided. ADR 0029's inner-loop argument holds with copy-on-write and did not without it. The consistency matters as much as the speed: three consecutive snapshots took 0.13, 0.12 and 0.13 seconds, against a dir second snapshot that never finished. One honest counter-observation recorded: launching onto the fresh copy-on-write pool was slower, 20.2s against 14.3s, because the image had to be unpacked into a pool that had never seen it. Paid once per pool, and dwarfed by what snapshotting saves, but it went the other way. Doing the measurement produced the answer to how the lab installs on a clean machine, because both failure modes appeared while doing it. Installed is not available: the daemon was present with units disabled and no group. Issue 007. Available is not adequate, and this is worse: with the storage tooling absent everything worked and snapshots were seventy-six times slower. Nothing failed, nothing warned. That is a variant the mesh has not catalogued — its usual failure is reported success and did nothing; this is reported success and did it seventy-six times slower, which no error surface catches because nothing is wrong. So the lab verifies CAPABILITY, never installation, and refuses to run degraded rather than warning — a warning about a slow inner loop is read once and ignored forever. Prerequisites may arrive from a mesh module or from the lab's own bootstrap, and the second path is required rather than convenient: a lab installable only by a mesh cannot host the development of the mesh that installs it. The lab is the second thing installed by hand, after the node host, and for the same reason: something has to be first, and pretending otherwise produces a circularity papered over by a script nobody exercises. |
||
|
|
98bcd5cc49 |
Measure the lab's inner loop — it is too slow, for a fixable reason
The lifecycle design closed on a question that was measurable rather than arguable, so it was measured. One virtual machine on a workstation with hardware virtualisation and NVMe. Raising: the launch call returns in 3.4s, the machine is actually usable after 14.3s. The gap is a design constraint — raise must wait for the second number, because reporting the first would be transport reported as effect, which is the mesh's own recurring failure. Snapshot: 9.9s and 1.6 GB for a 1.5 GB instance. A dir snapshot is a full copy; nothing is shared. Restore: 10.4s, usable again after 20.1s. The second snapshot exceeded two minutes and never completed. That is the more troubling number: snapshot cost here is not merely high, it is unpredictable, and a loop with a variable multi-minute step is one nobody trusts. Projected to a four-machine scenario, a reset-and-rerun cycle is about a minute and a half at best and unbounded at worst, before any of the mesh's own work begins. That is too slow for an inner loop, and ADR 0029's whole argument — that making the bootstrap path the inner loop turns the least-exercised code into the most-exercised — holds only while resetting is cheap. The cause is not virtual machines. Hardware virtualisation is present and machines boot in fourteen seconds. It is that the daemon offers exactly one storage driver, dir, which has no copy-on-write and therefore no cheap snapshot. The btrfs kernel module is available; btrfs-progs is simply not installed, which is the entire reason the driver is absent. The copy-on-write comparison was deliberately NOT run, because running it would mean installing a package by hand — which the rules forbid and which would have made the measurement unreproducible. So the honest statement is that the current configuration is too slow and the likely fix is known but unverified, rather than that btrfs fixes it. |
||
|
|
a253afe020 |
Scenario lifecycle, and how two scenarios coexist
ADR 0032: a scenario is a closed address space. Every segment materialises as its own isolated link belonging to one instance, so two scenarios raised from the same declaration hold the same addresses and never meet. The declaration keeps its literal addresses and they mean what they say — allocating from a pool would have made them a fiction, so a scenario reproducing a specific topology would stop reproducing it. The constraint that follows shapes everything: the lab never reaches into a scenario over IP. It talks to machines through the virtualisation layer's own channel. If it reached them by address, the workstation would need a route into each scenario, and two carrying the same prefix would give it two routes to one destination — failing not with an error but by one scenario's traffic arriving in another. That also makes reachability an honest question. Can this machine reach that one is asked from INSIDE, by executing on the first, rather than probed from a workstation that is not on the network and whose opinion would be a different question with a misleadingly similar answer. The lifecycle itself: six verbs, of which raise and destroy are enough to be useful and the rest are what make repetition cheap. Raising is convergent rather than incremental, because a lab behaving differently from the thing it tests teaches the wrong habit. A failed raise leaves the wreckage standing. Tearing down on failure destroys the only evidence, which is backwards — a scenario that failed to raise is more interesting than one that succeeded. Snapshots are whole-scenario. Per-machine would be cheaper and wrong: the mesh keeps state spanning nodes, so restoring one machine while its peers move on produces a mesh that has never existed, and faults found there would be artefacts of the lab. Closes the declaration's open question about running several scenarios at once. |
||
|
|
e88a6df924 |
Public networks are unrelated, and routed rather than bridged
Caught in review: every public address sat in one /24, which made the three of them look like one network. They are not. The internet is a very large number of unrelated networks routing to each other, and a machine in one is many hops from a machine in another with no shared broadcast domain between them. Putting them in one prefix would have quietly made four false things true in the lab: machines resolving each other by ARP and talking directly, TTL never decrementing, broadcast and multicast crossing between them, and any two being adjacent. The third is not hypothetical. The as-is layer records that mesh names are deliberately not multicast names, after a delay and a one-node-only failure mode. A lab where the internet is one broadcast domain would let a node discover a peer by multicast that it could never discover in production, and report success — the exact false green this effort exists to prevent. So a scenario has one public segment per public NETWORK, each with its own unrelated prefix, wired together through a router and never onto a shared bridge. That is a property of how the lab wires them rather than a field anyone sets, because no correct scenario has two public networks adjacent. Addresses now spread across all three RFC 5737 ranges plus RFC 3849 /48s, chosen to look nothing like each other, and a foreign private network uses someone else's RFC 1918 range rather than a documentation one. All three examples in the document rewritten, since two of them still showed a single flat internet segment and contradicted the new rule. |
||
|
|
a873088140 |
A worked example: the whole model applied to an ordinary mesh
The shape research 004 identified — one machine with a routable address, one publicly named but behind a household connection, one stationary on that network, one that roams — written out with every field the model has, in role names and documentation addresses. It shows the ISP modem doing nothing, because in bridge mode it is a media converter: it changes the physical medium and leaves the packets alone, so it creates no IP-level fact and appears nowhere. In router mode it would be a second gateway and publishing would need a rule on both, which is the one case the model still cannot express. It shows two segments sharing one gateway declaration, which means one gateway machine, and a policy rule between them that is asymmetric because useful ones almost always are. And it shows what is deliberately absent. Research 004 recorded overlay addresses, hub election and names for exactly this topology, and none of them appear: a scenario must not state what the mesh is responsible for. Given the declaration, whether a hub is elected, whether the NATed machine's endpoint is learned, and whether the roaming machine re-forms after moving are all observed rather than arranged. The absence is the point. The run at the end moves one identity through four positions — home, foreign network, asleep, home again — against a foreign gateway whose mapping expires in 30 seconds, which is why a number is there rather than a boolean. |
||
|
|
b1874f1d0f |
Segment policy, shared gateways, and what the model leaves out
Asked whether a real setup is coverable — router, modem, access points — the answer splits, and one part was a genuine gap. Most equipment is invisible and the omission is deliberate. The test: does the device change what an IP packet can do? A switch moves frames within a segment. An access point bridges wireless clients onto one — a machine on wifi and a machine on cable are the same machine to IP. A controller configures equipment and has no packets of its own. Modelling any of them adds a fixture with no fault to catch. Two entries in that list do matter. A modem in bridge mode is a media converter and invisible; in router mode it is a second gateway, which is double NAT — expressible as nested segments, but publishing through two gateways still is not, and that is now named as the one real absence. And VLANs are segments, which exposed the gap: inter-segment policy was inexpressible. inbound: is a HOST firewall, per machine. A segmented router enforcing rules between networks is a different thing and blocks traffic regardless of what the destination thinks — a node behind such a rule cannot be reached even by a peer that knows exactly where it is. policy: states it as a fact about a pair rather than a property of either, defaulting to allowed and asymmetric by design, because the useful configuration is almost always one-directional. Segments may also share a gateway: identical gateway declarations mean one gateway machine, not two, because that is what a VLAN-capable router is — and two routers sharing an address would not work anyway. |
||
|
|
274bd3b304 |
Close the missing axes — and address family changes the model
Address family was not a field. IPv6 usually has no NAT, so a machine behind a household gateway is typically unforwardable on v4 and DIRECTLY ATTACHED on v6, at the same moment. The three positions therefore apply per family, and reachability is a property of (machine, family) rather than of a machine. The consequence is bigger than the syntax: 'can these two nodes reach each other' stops being a yes/no question. It is asked once per family, and the asymmetric answers are the interesting ones. A mesh treating reachability as one fact per node reaches a peer over one family, fails over the other, and reports whichever it tried. That distinction did not exist in the model and would have been found by a failure rather than by reading. Two fields follow from it. inbound: allow|deny became necessary because with NAT unreachability was implied by topology, while a globally routable v6 address is reachable unless something refuses — so refusing has to be sayable or v6 addressing silently implies reachability. And nat: became a list of families rather than a boolean, because a real gateway translates v4 and routes v6 and a boolean cannot say that. mapping_ttl closes the keepalive gap: a mesh holding a connection through NAT without refreshing it works perfectly until the far side goes quiet for longer than the mapping lives. segments[].mtu closes the fragmentation gap: an overlay adds a header, so a tunnel over a reduced-MTU path establishes a connection and then silently drops large packets. at: takes a list, so a multi-homed machine is expressible — which the model already implicitly required, since a border machine sits on two segments. v6 uses RFC 3849 documentation space, the exact counterpart of the RFC 5737 rule and load-bearing for the same reason. Remaining: nested forwarding and an address changing in place, both extensible when needed. Path quality stays deliberately out — it changes performance, not correctness, and modelling it makes a network simulator rather than a fixture. |
||
|
|
b944904f1a |
Audit the scenario model for generality, and fix what it found
The question is not whether the model covers our mesh but whether it can express any mesh. Audited against the axes a deployment varies along, with the standard being every property that changes how the mesh BEHAVES rather than every property a network has — bandwidth does not change correctness, MTU does. One real bug, now fixed. A segment with no gateway was read as the internet, which made an isolated network inexpressible: a LAN with no route out would have been treated as public and forced onto documentation addresses. Segments now state kind: public or private, and a private segment with no gateway is an island. A mesh spanning a site with no internet is a real topology. One modelling error, now corrected. The three positions were framed by ownership — a gateway you control versus one you do not. The axis is forwardability. Carrier-grade NAT is your own connection and is still unforwardable, so it belongs with the café network. Gateways gain forwardable:, independent of nat:, and publishing through an unforwardable one is a declaration error because that is the constraint being reproduced. Three genuine gaps recorded in priority order. Address family: cidr is implicitly v4, and a v6-only node is not exotic — a mesh that assumes v4 fails there completely rather than partially, which makes this a second world rather than a refinement. Expiring NAT mappings: without them keepalive behaviour is hoped for rather than tested, and for a mesh mostly behind NAT that is the fault that shows up after an idle night. MTU: tunnels fragment, and a smaller-MTU path establishes a connection that then silently drops large packets — the exact shape this effort exists to stop shipping. Latency and loss are deliberately out: they change performance, not correctness, and modelling them makes a network simulator rather than a fixture. Also adds a NAT primer, because the three positions are consequences of it and the document should not assume the reader already knows why a mesh dials outward and never inward. |
||
|
|
e65e5809dc |
The scenario declaration gets a real network model
forwarded: [443] was the tell. It implied a destination-NAT rule while never saying from which address, and the address is the whole point: a household's public address is what a peer records as the endpoint when a machine there dials out, and what a public name for a published machine there resolves to. It was decoration in the old shape and is load-bearing in this one. The model now names three positions a machine can be in, because they are genuinely different and the mesh has to cope with all three. Directly attached, with its own routable address. Behind a gateway you control, reachable only through a forwarded port at the gateway's address. Behind a gateway you do not control, reachable not at all, with an apparent address belonging to someone else's router that changes when the machine moves. The third is the hard one and the one that breaks reachability assumptions first. A gateway now carries three facts instead of a boolean: the parent segment, the address the world sees the network as, and whether addresses are translated — so a routed range is expressible as well as ordinary household NAT. published names the gateway it forwards through, which is how a machine on a LAN that itself has a public address is stated, and publishing on a foreign gateway is a declaration error because that is exactly the constraint being reproduced. Moving a machine between positions becomes a lifecycle operation rather than a declaration: the same identity at home, then on a foreign network, then asleep, in one run. Whether the overlay survives that and notices the endpoint changed is observed, never arranged. All three RFC 5737 ranges are now allocated a job — the internet segment, a foreign network, and a spare — with private segments kept byte-identical to production because those addresses mean the same everywhere. New open question worth having: a real gateway forgets NAT mappings after a timeout, and whether a scenario can say so decides whether keepalive behaviour is testable or merely hoped for. |