The scenario model, the lifecycle, and what the lab actually costs #6
@@ -0,0 +1,60 @@
|
||||
---
|
||||
status: active
|
||||
initiated: 2026-08-24
|
||||
touches:
|
||||
- 03-DESIGN/01-to-be/03-scenario-lifecycle.md
|
||||
- 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
|
||||
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
|
||||
became: []
|
||||
---
|
||||
|
||||
# 010 — What the lab's inner loop actually costs
|
||||
|
||||
## What is being investigated
|
||||
|
||||
[The lifecycle design](../../03-DESIGN/01-to-be/03-scenario-lifecycle.md) closes on an open
|
||||
question that is measurable rather than arguable:
|
||||
|
||||
> **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the
|
||||
> inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop
|
||||
> is slow, and everything above is theory.
|
||||
|
||||
Measured on a workstation, 2026-08-24. Numbers in [`measurements.md`](measurements.md).
|
||||
|
||||
## Why it matters
|
||||
|
||||
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) makes the
|
||||
bootstrap scenario the inner development loop for tiers 0 and 1 — the argument being that
|
||||
raising a node from nothing stops being the least-exercised path and becomes the most-exercised
|
||||
one. **That argument is only true if raising and resetting are cheap.** A loop that costs
|
||||
minutes is a loop people avoid, and the least-exercised path stays least-exercised.
|
||||
|
||||
## Status
|
||||
|
||||
Measured, and the finding is a blocker rather than a data point.
|
||||
|
||||
**A snapshot on the current host is a full copy of the machine's disk.** 1.6 GB and ten seconds
|
||||
for one small virtual machine at best — and over two minutes when observed a second time. Cost
|
||||
scales with the number of machines and the size of their disks, not with what changed.
|
||||
|
||||
**With copy-on-write it is 0.13 seconds and costs the delta.** Verified, not assumed. The
|
||||
projected four-machine reset cycle falls from roughly ninety seconds to roughly fifteen, of
|
||||
which almost all is a boot that cannot be avoided.
|
||||
|
||||
The cause is not virtual machines and not incus. It is that the host offers incus exactly one
|
||||
storage driver, `dir`, which has no copy-on-write and therefore no cheap snapshot. The kernel
|
||||
supports btrfs; the userspace tool that would let incus use it is simply not installed.
|
||||
|
||||
So the question *"is the lab's inner loop fast enough"* currently answers itself the wrong way,
|
||||
for a reason that is one declared package away from being fixed — and declaring packages is
|
||||
something the mesh already does.
|
||||
|
||||
## Open questions
|
||||
|
||||
| Question | Why it matters |
|
||||
|---|---|
|
||||
| ~~How much does a copy-on-write pool actually improve it?~~ **Measured: snapshot 9.9 s → 0.13 s, restore 10.4 s → 0.80 s, three snapshots sharing 1.36 GB rather than costing 4.8 GB.** The projected four-machine cycle falls from ~90 s to ~15 s. | Answered. The inner-loop argument holds *with* copy-on-write and did not without it. |
|
||||
| How does the lab install its own prerequisites on a clean machine? | The lab needs a virtualisation daemon, copy-on-write tooling and a pool before it can do anything — and it cannot depend on the mesh for them, since it is where the mesh is built. |
|
||||
| Why was the second snapshot more than twelve times slower than the first? | If snapshot cost is unpredictable rather than merely high, that is worse — a loop with a variable multi-minute step is one nobody trusts. |
|
||||
| Does a scenario snapshot need the machines stopped? | Stateless snapshots of a running virtual machine capture the disk but not memory. Whether a mesh restored that way is coherent is not established. |
|
||||
| What is the cost at scenario scale — four machines rather than one? | Only single-machine numbers were taken. If the operation is serial, four machines is four times the wait. |
|
||||
@@ -0,0 +1,141 @@
|
||||
---
|
||||
effort: 010-lab-inner-loop-cost
|
||||
updated: 2026-08-24
|
||||
---
|
||||
|
||||
# Measurements
|
||||
|
||||
Taken 2026-08-24 on a workstation with hardware virtualisation available, an NVMe-backed ext4
|
||||
root, and 300 GB free. One virtual machine, 1 GiB memory, 2 CPUs, from a cached distribution
|
||||
image.
|
||||
|
||||
## The environment, before anything ran
|
||||
|
||||
| Fact | Value | Consequence |
|
||||
|---|---|---|
|
||||
| Hardware virtualisation | present | virtual machines run at native speed; the choice in [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md) is not paying an emulation penalty |
|
||||
| Storage drivers the daemon offers | **`dir` only** | no copy-on-write, therefore no cheap snapshot |
|
||||
| Host filesystems | ext4 throughout | nothing copy-on-write to put a pool on |
|
||||
| btrfs kernel module | **available** | the kernel can do it |
|
||||
| `btrfs-progs` | **not installed** | which is the entire reason the driver is absent |
|
||||
|
||||
The last two rows are the finding. The daemon advertises only `dir` because the userspace tool
|
||||
for anything better is missing — not because the host cannot do better.
|
||||
|
||||
## Raising a machine
|
||||
|
||||
| Step | Time |
|
||||
|---|---|
|
||||
| launch call returns | **3.4 s** |
|
||||
| machine actually usable — a command executes on it | **14.3 s** |
|
||||
|
||||
The gap matters for the lifecycle design: `raise` returning is not the same as the scenario
|
||||
being ready, so the verb has to wait for the second number, not report the first. Reporting
|
||||
the first would be the mesh's own recurring failure — transport reported as effect.
|
||||
|
||||
## Snapshot and restore
|
||||
|
||||
| Operation | Time | Disk |
|
||||
|---|---|---|
|
||||
| snapshot, first | **9.9 s** | **+1.6 GB** |
|
||||
| snapshot, second | **> 120 s — did not complete** | — |
|
||||
| restore call returns | **10.4 s** | — |
|
||||
| machine usable again | **20.1 s** total | — |
|
||||
|
||||
Instance on disk before snapshotting: 1.5 GB. Snapshot directory afterwards: 1.6 GB. **A `dir`
|
||||
snapshot is a full copy** — the storage cost equals the instance, and nothing is shared.
|
||||
|
||||
Implied copy throughput on the first snapshot is roughly 160 MB/s, which is far below what the
|
||||
underlying NVMe can do and is consistent with a real, durable copy rather than a metadata
|
||||
operation.
|
||||
|
||||
**The second snapshot is the more troubling number.** It exceeded two minutes and was still
|
||||
running when the observation was cut off; only the first snapshot exists. Whatever the cause —
|
||||
page cache exhausted by the preceding restore, writeback contention — the practical
|
||||
consequence is that snapshot cost here is **not merely high, it is unpredictable**.
|
||||
|
||||
## What this projects to
|
||||
|
||||
A four-machine scenario, taking the optimistic single-machine numbers and assuming the
|
||||
operations are serial:
|
||||
|
||||
| | one machine | four machines |
|
||||
|---|---|---|
|
||||
| raise, to usable | 14 s | ~57 s |
|
||||
| snapshot | 10 s, 1.6 GB | ~40 s, 6.4 GB |
|
||||
| restore, to usable | 20 s | ~80 s |
|
||||
|
||||
A reset-and-rerun cycle is therefore **around a minute and a half at best**, and unbounded at
|
||||
worst, before any of the mesh's own work begins.
|
||||
|
||||
## The judgement
|
||||
|
||||
**This is too slow for an inner loop**, and the reason is not the design.
|
||||
|
||||
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) argues that
|
||||
making the bootstrap path the inner development loop turns the least-exercised code in the
|
||||
system into the most-exercised. That argument holds only while resetting is cheap. At a minute
|
||||
and a half a cycle, with occasional multi-minute stalls, the loop is one a person works around
|
||||
— and the path stays under-exercised for exactly the reason it always was.
|
||||
|
||||
Nothing about virtual machines causes this. Hardware virtualisation is present and the machines
|
||||
boot in fourteen seconds. **The cost is entirely the storage driver**, and the driver is absent
|
||||
because one userspace package is not installed on the host.
|
||||
|
||||
The mesh already has the mechanism for that: a module declares a package, and a hook makes it a
|
||||
working capability — which is precisely what was just done for the virtualisation daemon
|
||||
itself, and what [`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)
|
||||
is about.
|
||||
|
||||
## The fix, measured
|
||||
|
||||
The comparison was subsequently run. One package — `btrfs-progs`, no dependencies — installed by
|
||||
hand, the daemon restarted so it re-detected drivers, a copy-on-write pool created on a loop
|
||||
file, and the identical image launched onto it.
|
||||
|
||||
| Operation | `dir` | copy-on-write | |
|
||||
|---|---|---|---|
|
||||
| snapshot | 9.9 s, then **> 120 s** | **0.13 s** | ~76× faster, and *consistent* |
|
||||
| snapshot again | — | 0.12 s | |
|
||||
| snapshot a third time | — | 0.13 s | |
|
||||
| restore call | 10.4 s | **0.80 s** | ~13× faster |
|
||||
| restore, to usable | 20.1 s | **10.5 s** | the remainder is boot, which is irreducible |
|
||||
| three snapshots, storage | ~4.8 GB | **1.36 GB total, shared** | cost is the delta, not the disk |
|
||||
|
||||
**The fix is real, and larger than expected.** Snapshot goes from ten seconds to a tenth of a
|
||||
second, and — more importantly — from *wildly variable* to *flat*. Three consecutive snapshots
|
||||
took 0.13, 0.12 and 0.13 seconds. On `dir` the second snapshot never finished.
|
||||
|
||||
Storage stops scaling with the machine and starts scaling with what changed: three snapshots of
|
||||
a 1.5 GB instance occupied 1.36 GB in total, because they share.
|
||||
|
||||
### What it projects to
|
||||
|
||||
A four-machine reset-and-rerun cycle, the operation the inner loop repeats most:
|
||||
|
||||
| | `dir` | copy-on-write |
|
||||
|---|---|---|
|
||||
| snapshot the scenario | ~40 s, 6.4 GB | **~0.5 s**, delta-sized |
|
||||
| restore it | ~40 s + boot | **~3 s** + boot |
|
||||
| **cycle** | **~90 s, unbounded at worst** | **~15 s, dominated by boot** |
|
||||
|
||||
At fifteen seconds, dominated by a boot that cannot be avoided, the inner loop is viable and
|
||||
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)'s argument holds.
|
||||
At ninety it did not.
|
||||
|
||||
### One honest counter-observation
|
||||
|
||||
Launching onto the fresh copy-on-write pool was **slower** — 20.2 s to usable against 14.3 s —
|
||||
because the image had to be unpacked into a pool that had never seen it. That cost is paid once
|
||||
per pool, not per scenario, and it is dwarfed by what snapshotting saves. But it is a real
|
||||
number and it went the other way.
|
||||
|
||||
### State this left behind
|
||||
|
||||
Recorded because hand-made state is exactly what the mesh's rules exist to prevent, and it must
|
||||
be declared properly rather than left as an artefact of a measurement:
|
||||
|
||||
- `btrfs-progs` installed by hand. Its installation regenerated the boot initramfs, a side
|
||||
effect worth knowing about.
|
||||
- The daemon restarted once, to re-detect drivers.
|
||||
- The test pool and instance were **removed**; the pool the lab actually needs does not exist.
|
||||
@@ -0,0 +1,71 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-08-23
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
---
|
||||
|
||||
# 32. A scenario is an isolated address space, and the lab never reaches into it over IP
|
||||
|
||||
## Context
|
||||
|
||||
A scenario declares literal addresses —
|
||||
[the declaration](../03-DESIGN/01-to-be/02-scenario-declaration.md) is full of them, and it has
|
||||
to be, because reproducing *published but behind NAT* means saying which address the world sees.
|
||||
|
||||
That raises a question the declaration left open: **two scenarios at once.** Several agents
|
||||
working means several scenarios, and the lab design already calls that a requirement. But two
|
||||
scenarios built from the same declaration want the same addresses, and there are only three
|
||||
documentation ranges in existence.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Allocate addresses from a pool at raise time**, rewriting the declaration's literals.
|
||||
Rejected. It makes the addresses in a declaration a fiction, so a scenario reproducing a
|
||||
specific topology no longer reproduces it; it breaks the RFC-range validation, since
|
||||
allocated addresses would have to come from somewhere real; and the numbers a person reads
|
||||
in the file stop being the numbers they will see in a capture.
|
||||
2. **One scenario at a time.** Rejected — it is the requirement, not an inconvenience. A gate
|
||||
an agent has to queue for is a gate that gets bypassed.
|
||||
3. **Give each scenario its own network stack, so the addresses do not collide.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**A scenario is a closed address space.** Every segment materialises as its own isolated link,
|
||||
belonging to one scenario instance. Two scenarios raised from the same declaration hold the same
|
||||
addresses and never meet, because nothing joins their links.
|
||||
|
||||
The declaration therefore keeps its literal addresses, and they mean exactly what they say.
|
||||
|
||||
**The consequence that constrains everything else: the lab never reaches into a scenario over
|
||||
IP.** It talks to a machine through the virtualisation layer's own channel — the same way one
|
||||
executes a command in a container without the container being routable.
|
||||
|
||||
That is not a preference. If the lab reached machines by address, the workstation running it
|
||||
would need a route into each scenario, and two scenarios carrying the same prefix would give it
|
||||
two routes to the same destination. Concurrency would be impossible, and it would fail in the
|
||||
worst available way: not with an error, but by one scenario's traffic arriving in another.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Scenarios are concurrent by construction, with no allocation, no bookkeeping and no limit
|
||||
beyond the machine's capacity.
|
||||
- The three documentation ranges stop being a scarce resource. Every scenario may use all of
|
||||
them, because no two scenarios share a link.
|
||||
- **The lab cannot use IP to check anything**, which is more of a constraint than it first
|
||||
appears: *"can this machine reach that one"* has to be asked **from inside the scenario**, by
|
||||
executing on a machine, rather than probed from outside. That is the honest way to ask it
|
||||
anyway — reachability from the workstation is not the question.
|
||||
- A scenario is a unit that can be paused, snapshotted and destroyed whole, because nothing
|
||||
outside holds a reference into it.
|
||||
- The lab needs a scenario **instance** identity distinct from the scenario name in the
|
||||
declaration: the declaration is a kind, and several instances of one kind may exist.
|
||||
- **The workstation is not on the scenario's network, so it is not a node in it.** Anything a
|
||||
developer wants to reach — a web interface, a database — needs an explicit, deliberate
|
||||
forward out of the scenario, which is a feature rather than a gap: nothing leaks by default.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0031](0031-the-lab-provides-the-underlay.md) — the declaration whose literal addresses
|
||||
this preserves.
|
||||
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the lifecycle jobs this shapes.
|
||||
@@ -17,31 +17,191 @@ everything in the lab hangs off, so it is worth getting small.
|
||||
It states what a hosting provider and a home router would provide, and nothing the mesh is
|
||||
responsible for ([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)).
|
||||
|
||||
## Public networks are unrelated, and routed rather than bridged
|
||||
|
||||
The internet is not a network. It is a very large number of unrelated networks that route to
|
||||
each other, and a machine in one is **many hops** from a machine in another with no shared
|
||||
broadcast domain between them.
|
||||
|
||||
So a scenario does not have *an* internet segment. It has **one public segment per public
|
||||
network**, each with its own unrelated prefix, and the lab wires them together **through a
|
||||
router, never onto a shared bridge**.
|
||||
|
||||
That distinction is load-bearing, and putting several public addresses in one prefix would
|
||||
quietly make four things true that are false in reality:
|
||||
|
||||
| If public addresses share a segment | Reality |
|
||||
|---|---|
|
||||
| machines resolve each other by ARP and talk directly | they are routed, many hops apart |
|
||||
| TTL never decrements | every hop decrements it |
|
||||
| broadcast and multicast reach across | neither crosses a router |
|
||||
| any two are adjacent | adjacency is the exception, not the rule |
|
||||
|
||||
The third is not hypothetical here. The mesh has already been bitten by multicast name
|
||||
resolution — [`00-as-is/01-mesh-and-transport.md`](../00-as-is/01-mesh-and-transport.md)
|
||||
records that mesh names are deliberately not multicast names, after a delay and a
|
||||
one-node-only failure mode. A lab where "the internet" is one broadcast domain would let a node
|
||||
discover a peer by multicast that it could never discover in production, and report success.
|
||||
|
||||
**The rule: `kind: public` segments are routed to one another and never bridged.** It is a
|
||||
property of how the lab wires them, not a field anyone sets, because there is no correct
|
||||
scenario in which two public networks are adjacent.
|
||||
|
||||
### Which addresses to use
|
||||
|
||||
RFC 5737 reserves three ranges and RFC 3849 reserves one v6 prefix. Each public network takes
|
||||
its own, and they are chosen to look nothing like each other — because in reality they would
|
||||
not:
|
||||
|
||||
| Public network | v4 | v6 |
|
||||
|---|---|---|
|
||||
| a hosting provider | `192.0.2.0/24` | `2001:db8:a::/48` |
|
||||
| a household ISP | `198.51.100.0/24` | `2001:db8:b::/48` |
|
||||
| a mobile or foreign network | `203.0.113.0/24` | `2001:db8:c::/48` |
|
||||
|
||||
Three is enough for the topologies that matter, and a fourth public network can subnet one of
|
||||
them — an ISP handing out `198.51.100.0/25` and `198.51.100.128/25` to two customers is exactly
|
||||
what really happens.
|
||||
|
||||
Private segments still use RFC 1918 and stay byte-identical to production.
|
||||
|
||||
## What NAT does, and why the design turns on it
|
||||
|
||||
A household or office has **one** address the outside world can see, and **many** machines
|
||||
behind it. Network address translation is what reconciles those.
|
||||
|
||||
When a machine inside dials out, the gateway rewrites the packet's source from the private
|
||||
address to the public one, **remembers the mapping**, and rewrites the replies on the way back.
|
||||
Four consequences follow, and every one of them shapes this design:
|
||||
|
||||
1. **Outbound works; inbound does not.** A mapping exists only because something inside started
|
||||
a conversation. Nothing outside can start one — there is no mapping to look up, and no way
|
||||
to know which internal machine was meant.
|
||||
2. **A forwarded port is a permanent mapping made by hand**, in the inbound direction:
|
||||
*anything arriving at the public address on 443 goes to this machine.* That is the only way
|
||||
a machine behind NAT becomes reachable, and it requires control of the gateway.
|
||||
3. **Mappings expire.** A gateway forgets one that goes unused. This is why anything holding a
|
||||
connection through NAT sends keepalives, and why a mesh that does not is fine until it is
|
||||
idle.
|
||||
4. **From outside, every machine behind the gateway looks like one address.** Identity and
|
||||
address stop corresponding.
|
||||
|
||||
This is why the mesh dials outward and never inward
|
||||
([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md)), why a hub exists at
|
||||
all, and why a node's endpoint is something a peer **learns** from arriving packets rather than
|
||||
something anyone configures.
|
||||
|
||||
**Carrier-grade NAT** is the same mechanism applied by an ISP: your own gateway gets a private
|
||||
address too, and the public one is shared with strangers. Nothing can be forwarded, because the
|
||||
rule would have to live on equipment you do not own. Common on mobile connections and
|
||||
increasingly on fixed ones.
|
||||
|
||||
## Three positions a machine can be in
|
||||
|
||||
The underlay's whole job is to reproduce **where a machine sits relative to the internet**,
|
||||
because that is what the mesh has to cope with and what only production currently exercises.
|
||||
There are three positions, and they are genuinely different:
|
||||
|
||||
| Position | Reachable from outside | Apparent address | Example |
|
||||
|---|---|---|---|
|
||||
| **Attached** | yes, at its own address | its own | a hosted server |
|
||||
| **Behind a forwardable gateway** | only through a forwarded port, at the *gateway's* address | the gateway's | a machine at home |
|
||||
| **Behind an unforwardable gateway** | **no** | someone else's, and it changes | a laptop on a café network; anything behind carrier-grade NAT |
|
||||
|
||||
The axis is **forwardability, not ownership** — which is worth stating because the obvious
|
||||
framing gets it wrong. Carrier-grade NAT is *your* connection and is still unforwardable, so it
|
||||
belongs in the third row alongside the café. What the mesh has to cope with is whether an
|
||||
inbound mapping can be made, not who owns the equipment.
|
||||
|
||||
The third position is the hard one. A machine there can dial out and nothing more: it cannot be
|
||||
published, its apparent address belongs to a router it does not control, and that address
|
||||
changes when it moves. Every assumption a mesh makes about reachability breaks there first.
|
||||
|
||||
A declaration has to be able to say all three, and to move a machine between them.
|
||||
|
||||
## Reachability is per address family, not per machine
|
||||
|
||||
Adding IPv6 is not a field. It changes the position model, and the reason is worth stating
|
||||
before the syntax.
|
||||
|
||||
**IPv6 usually has no NAT.** A machine behind a household gateway can hold a *globally routable*
|
||||
v6 address while its v4 address is private and unforwardable. The same machine, at the same
|
||||
moment, is in **two different positions at once**:
|
||||
|
||||
| | IPv4 | IPv6 |
|
||||
|---|---|---|
|
||||
| a typical machine at home | behind an unforwardable-or-forwardable gateway | **attached**, directly reachable |
|
||||
| a machine on mobile data | behind carrier-grade NAT | often attached, sometimes absent entirely |
|
||||
| a machine on an older network | attached or behind NAT | **no address at all** |
|
||||
|
||||
So the three positions apply **per family**, and a machine's reachability is a property of
|
||||
*(machine, family)* rather than of the machine. A mesh that treats reachability as one fact per
|
||||
node will reach a peer over one family, fail over the other, and report whichever it tried.
|
||||
|
||||
That has a direct consequence for what the lab is for: *"can these two nodes reach each other"*
|
||||
stops being a yes/no question. It is asked once per family, and the interesting answers are the
|
||||
asymmetric ones.
|
||||
|
||||
The v6 documentation prefix is `2001:db8::/32` (RFC 3849) — the exact counterpart of the RFC
|
||||
5737 rule, and load-bearing for the same reason.
|
||||
|
||||
## The shape
|
||||
|
||||
```yaml
|
||||
scenario: published-behind-nat
|
||||
scenario: roaming-and-published
|
||||
|
||||
segments:
|
||||
wan:
|
||||
cidr: 203.0.113.0/24 # RFC 5737 — never routes on the real internet
|
||||
lan:
|
||||
cidr: 192.168.1.0/24
|
||||
behind: wan # NAT; the lab materialises a router
|
||||
hosting: # one public network
|
||||
kind: public
|
||||
cidr: [192.0.2.0/24, 2001:db8:a::/48]
|
||||
|
||||
isp-home: # another, unrelated — routed to it, not bridged
|
||||
kind: public
|
||||
cidr: [198.51.100.0/24, 2001:db8:b::/48]
|
||||
|
||||
home:
|
||||
kind: private
|
||||
cidr: [192.168.1.0/24, 2001:db8:b:1::/64]
|
||||
mtu: 1500
|
||||
gateway:
|
||||
to: isp-home
|
||||
address: [198.51.100.7, 2001:db8:b::7] # what the world sees this network as
|
||||
nat: [v4] # v4 is translated; v6 is routed, not translated
|
||||
forwardable: true
|
||||
mapping_ttl: 120s # an unused inbound mapping is forgotten after this
|
||||
|
||||
cafe: # a network we do not control
|
||||
kind: private
|
||||
cidr: [10.50.0.0/16] # v4 only — no v6 offered here at all
|
||||
mtu: 1400 # a tunnelled path, smaller than standard
|
||||
gateway:
|
||||
to: hosting # it reaches the world via a different public network
|
||||
address: [192.0.2.200]
|
||||
nat: [v4]
|
||||
forwardable: false # café wifi, or carrier-grade NAT
|
||||
mapping_ttl: 30s # aggressive, as carrier NAT tends to be
|
||||
|
||||
machines:
|
||||
anchor:
|
||||
segment: wan
|
||||
address: 203.0.113.10
|
||||
at: { segment: hosting, address: [192.0.2.10, 2001:db8:a::10] }
|
||||
|
||||
home-server:
|
||||
segment: lan
|
||||
address: 192.168.1.135
|
||||
forwarded: [443] # reachable from wan through the router
|
||||
at: { segment: home, address: [192.168.1.135, 2001:db8:b:1::135] }
|
||||
published:
|
||||
- { port: 443, on: home } # v4 only: 198.51.100.7:443 → 192.168.1.135:443
|
||||
inbound: allow # v6 is routable here, so this decides whether it is reachable
|
||||
|
||||
workstation:
|
||||
segment: lan
|
||||
address: 192.168.1.250
|
||||
at: { segment: home, address: [192.168.1.250, 2001:db8:b:1::250] }
|
||||
inbound: deny # a host firewall: dials out, accepts nothing
|
||||
|
||||
laptop:
|
||||
segment: detached # reachable by nothing until attached
|
||||
at: { segment: home, address: [192.168.1.98, 2001:db8:b:1::98] }
|
||||
|
||||
border: # a machine on two segments at once
|
||||
at:
|
||||
- { segment: home, address: [192.168.1.2] }
|
||||
- { segment: isp-home, address: [198.51.100.60] }
|
||||
|
||||
place:
|
||||
all: [host]
|
||||
@@ -50,41 +210,113 @@ place:
|
||||
snapshot: raised
|
||||
```
|
||||
|
||||
That is a complete bootstrap scenario. Nothing in it mentions the overlay, a hub, peering,
|
||||
names or certificates — all of which are outcomes to be observed.
|
||||
## What each part means, precisely
|
||||
|
||||
## The four parts
|
||||
**`segments`** — a broadcast domain with an address range, and a `kind:`. A segment is a single
|
||||
broadcast domain, which is precisely why several public networks cannot be one segment.
|
||||
|
||||
**`segments`** — the networks that exist. `behind:` declares NAT, and is the only place a
|
||||
router comes from: the lab materialises one without being asked, because NAT has to run
|
||||
somewhere. This is the one implicit machine in an otherwise explicit declaration.
|
||||
`kind: public` marks a segment that stands in for a public network — and there is normally more
|
||||
than one, unrelated to each other. `kind: private` is everything else. This is stated rather than inferred, and the earlier version inferred it — *a segment with
|
||||
no gateway is the internet* — which made an **isolated network inexpressible**: a LAN with no
|
||||
route out is a private segment with no gateway, and would have been read as public and forced
|
||||
to use documentation addresses. A mesh spanning a site with no internet access is a real
|
||||
topology, and the model has to be able to say it.
|
||||
|
||||
**`machines`** — what sits where. A machine has a segment and an address, and that is nearly
|
||||
all. `forwarded:` opens a port through the router, which is what makes *published but behind
|
||||
NAT* reproducible — the case that exists only in production today. `segment: detached` is a
|
||||
machine on no network, which is how a roaming node is expressed at rest.
|
||||
**`gateway:`** — how a segment reaches its parent, and this is where the previous version was
|
||||
too thin. It carries three facts, and all three are load-bearing:
|
||||
|
||||
**`place`** — what goes inside. `all:` applies to every machine; a machine name overrides for
|
||||
that machine. This is the only part that differs between the two scenario classes.
|
||||
- `to:` — the parent segment.
|
||||
- `address:` — **the address the outside world sees this network as.** For a household
|
||||
connection this is the public address the ISP hands out. It is not decoration: it is what a
|
||||
peer records as the endpoint when a machine here dials out, and what a public name for a
|
||||
published machine here resolves to.
|
||||
- `nat:` — **which families are translated**, as a list. `[v4]` is the ordinary modern case:
|
||||
v4 translated, v6 routed. `[v4, v6]` describes a gateway that translates both, which exists
|
||||
and is worth being able to reproduce. `[]` is a routed range, where machines keep their own
|
||||
addresses and the gateway only forwards.
|
||||
- `forwardable:` — whether an inbound mapping can be created. Independent of `nat:`, and the
|
||||
field that separates a home gateway from carrier-grade NAT. Publishing through a gateway with
|
||||
`forwardable: false` is a declaration error, because that is exactly the constraint being
|
||||
reproduced.
|
||||
- `mapping_ttl:` — how long an unused inbound mapping survives. This is what makes keepalive
|
||||
behaviour testable: a mesh that holds a connection through NAT without refreshing it works
|
||||
perfectly until the far side goes quiet for longer than this. Aggressive values reproduce
|
||||
carrier NAT; omitting it means mappings never expire, which no real gateway does.
|
||||
|
||||
**`snapshot`** — names the state once placement finishes, so a run can return to it without
|
||||
raising everything again. Snapshots are what make repetition cheap, and cheap repetition is
|
||||
what makes the bootstrap path the inner development loop rather than a ceremony.
|
||||
**`segments[].mtu`** — the largest packet the segment carries, defaulting to 1500. Lower values
|
||||
reproduce tunnelled and PPPoE paths. This matters because an overlay adds its own header: a
|
||||
tunnel over a 1400-byte path establishes a connection and then silently drops large packets,
|
||||
which is the shape of fault this whole effort exists to stop shipping.
|
||||
|
||||
**`machines[].inbound`** — `allow` or `deny`, a host firewall. Distinct from NAT and behaves
|
||||
differently: a machine can be perfectly routable and still refuse everything unsolicited, which
|
||||
is the normal state of a v6-addressed machine. Without this, v6 addressing would imply
|
||||
reachability, and it does not.
|
||||
|
||||
The lab materialises a machine to be the gateway. That is the one implicit machine in an
|
||||
otherwise explicit declaration, and it exists because NAT has to run somewhere.
|
||||
|
||||
**`machines[].at`** — segment and addresses, or a **list** of them for a machine on several
|
||||
segments at once. Multi-homing is not exotic: it is what a border machine is, and what any node
|
||||
with both a LAN and a WAN interface is. Each entry carries the addresses that machine holds on
|
||||
that segment, one per family.
|
||||
|
||||
Position follows from the pair, per family: on a `kind: public` segment a machine is attached;
|
||||
on a private one it is behind that segment's gateway, unless the gateway does not translate
|
||||
that family — in which case it is attached on that family and behind a gateway on the other.
|
||||
|
||||
**`machines[].published`** — a destination-NAT rule on a named gateway, stated as an outcome
|
||||
rather than a port list. `{ port: 443, on: home }` means the `home` gateway forwards its own
|
||||
`203.0.113.50:443` to this machine's `443`. The resulting public endpoint is derivable, which is
|
||||
the point: a scenario never writes an endpoint down, and the mesh has to discover it.
|
||||
|
||||
A machine may be published on **any gateway between it and a public network** — which is how *"our
|
||||
LAN also has a public IP"* is expressed, and why `on:` names the gateway rather than being
|
||||
implied. It cannot be published at all on a gateway the scenario models as foreign; attempting
|
||||
it is a declaration error, because that is precisely the constraint being reproduced.
|
||||
|
||||
**`at: detached`** — on no segment. A machine that exists and can reach nothing.
|
||||
|
||||
## Moving a machine is a lifecycle operation
|
||||
|
||||
`at:` states where a machine *starts*. Moving it is something a run does:
|
||||
|
||||
```
|
||||
move laptop → { segment: elsewhere, address: 198.51.100.23 }
|
||||
move laptop → detached
|
||||
move laptop → { segment: home, address: 192.168.1.98 }
|
||||
```
|
||||
|
||||
This is the roaming case made testable, and it is the one that finds the interesting faults.
|
||||
The same machine, the same identity, three positions in one run: at home where its peers can
|
||||
reach it directly, on a foreign network where it can only dial out and its apparent address
|
||||
belongs to a router it does not control, and asleep.
|
||||
|
||||
Whether the overlay survives that, re-forms, and is noticed to have changed endpoint is
|
||||
**observed**, never arranged
|
||||
([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)).
|
||||
|
||||
## Why the addresses are load-bearing
|
||||
|
||||
The routable segment uses RFC 5737 documentation space, and this is not a stylistic choice.
|
||||
The internet segment uses RFC 5737 documentation space, and this is not a stylistic choice.
|
||||
|
||||
The mesh decides *public versus private* by matching the address. A private range on the
|
||||
segment meant to be routable makes the hub test as unreachable, and **the mesh silently never
|
||||
forms** — no error, no failed step, just a mesh that does not exist. Research 004 calls this
|
||||
the single most important fact in its analysis.
|
||||
segment meant to be routable makes a would-be hub test as unreachable, and **the mesh silently
|
||||
never forms** — no error, no failed step, just a mesh that does not exist. Research 004 calls
|
||||
this the single most important fact in its analysis.
|
||||
|
||||
So the format should make this hard to get wrong rather than merely documented: a segment
|
||||
without `behind:` is a routable segment, and an address in it that is not documentation space
|
||||
is a declaration error, refused before anything is raised. That is
|
||||
Which range goes where is covered above, under *public networks are unrelated*: one range per
|
||||
public network, chosen to look nothing like each other.
|
||||
|
||||
Private segments use RFC 1918 and can be **byte-identical to production**, because those
|
||||
addresses mean the same thing everywhere. Only the public side is substituted, and only because
|
||||
it must be.
|
||||
|
||||
The format should make getting this wrong hard rather than merely documented: a segment without
|
||||
a `gateway:` is a public segment, and an address in it — including a gateway's `address:` — that
|
||||
is not documentation space is a declaration error, refused before anything is raised. That is
|
||||
[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) applied to a configuration
|
||||
file — the failure it prevents is silent, so the check has to be loud.
|
||||
file: the failure it prevents is silent, so the check has to be loud.
|
||||
|
||||
## The same declaration serves both classes
|
||||
|
||||
@@ -124,18 +356,284 @@ not first.
|
||||
- **Steps.** A scenario is a desired state. Anything expressed as an ordered list of actions
|
||||
belongs in the lifecycle, not the declaration.
|
||||
|
||||
## What the model deliberately does not contain
|
||||
|
||||
A real network is full of equipment: a modem, a router, switches, access points, controllers.
|
||||
Almost none of it appears here, and the omission is deliberate rather than an oversight.
|
||||
|
||||
**The test is whether a device changes what an IP packet can do.** If two machines can exchange
|
||||
packets, at the same addresses, with the same reachability and the same MTU, whether or not the
|
||||
device exists — then the device is invisible to the mesh, and modelling it would add a fixture
|
||||
with no fault to catch.
|
||||
|
||||
| Equipment | Modelled? | Why |
|
||||
|---|---|---|
|
||||
| **Switch** | no | Moves frames within a segment. Two machines on a switch are two machines on a segment. |
|
||||
| **Access point** | no | Bridges wireless clients onto a segment. A machine on wifi and a machine on cable are the same machine to IP. |
|
||||
| **Network controller** | no | Configures equipment. Its effects appear as segments and policy; it has no packets of its own. |
|
||||
| **Cabling, PoE, uplink speed** | no | Change performance and availability, not reachability. |
|
||||
| **Router / security gateway** | **yes — it *is* the gateway** | Translation, forwarding and inter-segment policy all live here. |
|
||||
| **ISP modem** | **only in router mode** | In bridge mode it is a media converter and invisible. Doing its own NAT, it is a second gateway — and that is double NAT. |
|
||||
| **VLANs** | **yes — they are segments** | Machines on separate VLANs cannot reach each other except through the router, which is the definition of a separate segment. |
|
||||
| **Inter-VLAN firewall rules** | **yes** — see below | A rule stopping one segment reaching another is a reachability fact, and the mesh will hit it. |
|
||||
|
||||
The two entries worth dwelling on are the ones where a common household setup produces
|
||||
something the mesh has to survive.
|
||||
|
||||
**A modem in router mode gives you double NAT.** Your gateway holds a private address from the
|
||||
modem, which holds the public one. Forwarding then requires a rule on *both*, and one of them
|
||||
may not be configurable. This is expressible as nested segments — `to:` chains — but publishing
|
||||
through two gateways is not, and remains open.
|
||||
|
||||
**Segmented networks are the common case, not the exotic one.** A router with separate networks
|
||||
for trusted machines, guests and devices is ordinary, and the rules between them are ordinary
|
||||
too. A mesh node on one segment and a mesh node on another are, as far as reachability goes, on
|
||||
different networks that happen to share a gateway.
|
||||
|
||||
## Segments may share a gateway
|
||||
|
||||
Several segments can name the same parent and the same external address. That is one router
|
||||
with several networks behind it, which is what a VLAN-capable gateway is:
|
||||
|
||||
```yaml
|
||||
segments:
|
||||
trusted:
|
||||
kind: private
|
||||
cidr: [192.168.1.0/24]
|
||||
gateway: { to: isp-home, address: [198.51.100.7], nat: [v4], forwardable: true }
|
||||
devices:
|
||||
kind: private
|
||||
cidr: [192.168.30.0/24]
|
||||
gateway: { to: isp-home, address: [198.51.100.7], nat: [v4], forwardable: true }
|
||||
```
|
||||
|
||||
Identical gateway declarations mean **one gateway machine**, not two. The lab materialises a
|
||||
single router serving both, because that is what the topology being reproduced is — and two
|
||||
routers sharing one address would not work anyway.
|
||||
|
||||
## Policy between segments
|
||||
|
||||
Sharing a gateway does not mean segments can reach each other. What they may do is stated
|
||||
separately, because it is a fact about a pair rather than a property of either:
|
||||
|
||||
```yaml
|
||||
policy:
|
||||
- { from: devices, to: trusted, allow: false } # devices may not initiate to trusted
|
||||
- { from: trusted, to: devices, allow: true } # the reverse is fine
|
||||
```
|
||||
|
||||
Default is `true` between segments behind the same gateway, matching a router with no rules
|
||||
configured. Asymmetry is the normal case and the reason this cannot be a single flag: the
|
||||
useful configuration is almost always one-directional.
|
||||
|
||||
This is **not** the same as `machines[].inbound`, and conflating them loses a real distinction:
|
||||
|
||||
| | Enforced by | Blocks |
|
||||
|---|---|---|
|
||||
| `policy` | the gateway, between segments | everything crossing, regardless of what the destination thinks |
|
||||
| `inbound` | the machine itself | unsolicited traffic that already reached it |
|
||||
|
||||
A mesh node behind a `policy` deny cannot be reached even by a peer that knows exactly where it
|
||||
is — and that is a real topology, not a contrived one.
|
||||
|
||||
## Worked example — a whole mesh of the ordinary kind
|
||||
|
||||
The shape research 004 identified: one machine with a routable address, one publicly named but
|
||||
behind a household connection, one stationary machine on that network, one that roams. Written
|
||||
out completely, with every field the model has.
|
||||
|
||||
```yaml
|
||||
scenario: the-ordinary-shape
|
||||
|
||||
segments:
|
||||
# ---- three unrelated public networks. Routed to each other, never bridged. ----
|
||||
|
||||
hosting: # where the always-on machine lives
|
||||
kind: public
|
||||
cidr: [192.0.2.0/24, 2001:db8:a::/48]
|
||||
|
||||
isp-home: # the household's uplink
|
||||
kind: public
|
||||
cidr: [198.51.100.0/24, 2001:db8:b::/48]
|
||||
|
||||
isp-mobile: # wherever the roaming machine happens to be
|
||||
kind: public
|
||||
cidr: [203.0.113.0/24, 2001:db8:c::/48]
|
||||
|
||||
# ---- private networks behind them ----
|
||||
|
||||
home: # the household network
|
||||
kind: private
|
||||
cidr: [192.168.1.0/24, 2001:db8:b:1::/64]
|
||||
mtu: 1492 # PPPoE on the uplink; 1500 if the line is not PPPoE
|
||||
gateway:
|
||||
to: isp-home
|
||||
address: [198.51.100.7, 2001:db8:b::7]
|
||||
nat: [v4] # v4 translated, v6 routed — the modern default
|
||||
forwardable: true # the household router is ours to configure
|
||||
mapping_ttl: 120s
|
||||
|
||||
devices: # optional: a segmented network on the same router
|
||||
kind: private
|
||||
cidr: [192.168.30.0/24]
|
||||
gateway:
|
||||
to: isp-home
|
||||
address: [198.51.100.7, 2001:db8:b::7] # identical → the SAME gateway machine
|
||||
nat: [v4]
|
||||
forwardable: true
|
||||
mapping_ttl: 120s
|
||||
|
||||
cafe: # a network we do not control
|
||||
kind: private
|
||||
cidr: [10.50.0.0/16] # RFC 1918 — someone else's private range
|
||||
mtu: 1400
|
||||
gateway:
|
||||
to: isp-mobile
|
||||
address: [203.0.113.129]
|
||||
nat: [v4]
|
||||
forwardable: false # carrier-grade NAT, or simply not ours
|
||||
mapping_ttl: 30s
|
||||
|
||||
policy:
|
||||
- { from: devices, to: home, allow: false }
|
||||
- { from: home, to: devices, allow: true }
|
||||
|
||||
machines:
|
||||
anchor: # routable, nothing in front of it
|
||||
at: { segment: hosting, address: [192.0.2.10, 2001:db8:a::10] }
|
||||
inbound: allow
|
||||
|
||||
home-server: # publicly named, behind the household connection
|
||||
at: { segment: home, address: [192.168.1.135, 2001:db8:b:1::135] }
|
||||
published:
|
||||
- { port: 443, on: home } # v4 reaches it only through the forward
|
||||
inbound: allow # and v6 reaches it directly, so this matters
|
||||
|
||||
workstation: # on the household network, not published
|
||||
at: { segment: home, address: [192.168.1.250, 2001:db8:b:1::250] }
|
||||
inbound: deny
|
||||
|
||||
laptop: # starts at home; moves during the run
|
||||
at: { segment: home, address: [192.168.1.98, 2001:db8:b:1::98] }
|
||||
inbound: deny
|
||||
|
||||
place:
|
||||
all: [host]
|
||||
anchor: [substrate]
|
||||
|
||||
snapshot: raised
|
||||
```
|
||||
|
||||
### What the ISP modem is doing here
|
||||
|
||||
**Nothing, if it is in bridge mode** — which is the ordinary arrangement when the household has
|
||||
its own router. A bridged modem is a media converter: it changes the physical medium and leaves
|
||||
the packets alone, so it creates no IP-level fact and appears nowhere above.
|
||||
|
||||
Were it in router mode it would be a second gateway, `home` would sit behind it rather than
|
||||
behind `internet` directly, and publishing would need a rule on both — the case the model
|
||||
cannot yet express.
|
||||
|
||||
### What is deliberately absent
|
||||
|
||||
Nothing here mentions overlay addresses, which node is the hub, who peers with whom, any name,
|
||||
or any certificate. Research 004 recorded all of those for this topology, and **a scenario must
|
||||
not state them** ([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)): they
|
||||
are what the mesh does, and a scenario that supplied them would be certifying its own work.
|
||||
|
||||
The absence is the point. Given the declaration above, whether a hub is elected, whether the
|
||||
NATed machine's endpoint is learned, whether the roaming machine re-forms after moving — all of
|
||||
it is observed.
|
||||
|
||||
### Running it
|
||||
|
||||
```
|
||||
move laptop → { segment: cafe, address: [10.50.3.23] } # it leaves the house
|
||||
move laptop → detached # it sleeps
|
||||
move laptop → { segment: home, address: [192.168.1.98, 2001:db8:b:1::98] }
|
||||
```
|
||||
|
||||
One identity, four positions, one run. The `mapping_ttl: 30s` on `cafe` means a connection held
|
||||
without refreshing dies while it is out — which is the point of putting a number there.
|
||||
|
||||
Note that the laptop's three positions are on **three different public networks**: at home it
|
||||
appears as `198.51.100.7`, at the café as `203.0.113.129`, and the machine it is trying to reach
|
||||
is on a fourth. None of them are adjacent, all of them are routed. That is the situation the
|
||||
mesh actually faces.
|
||||
|
||||
## Is this general? — the axes a setup can vary along
|
||||
|
||||
The question that matters is not *does this cover our mesh*, but **can it express any mesh**.
|
||||
|
||||
The standard applied is not "every property a network has". It is **every property that changes
|
||||
how the mesh behaves**. Bandwidth does not change correctness; MTU does.
|
||||
|
||||
| Axis | Values | Expressible | |
|
||||
|---|---|---|---|
|
||||
| **Reachability** | attached · forwardable gateway · unforwardable gateway · isolated | yes | the core of the model, and **per family** |
|
||||
| **Address family** | IPv4 · IPv6 · dual-stack · neither | yes | `cidr:` and `address:` take both; `nat:` names which families are translated |
|
||||
| **Interfaces per machine** | one · several | yes | `at:` takes a list |
|
||||
| **Reachability policy, host** | symmetric · asymmetric | yes | `inbound:` — a routable machine that refuses everything |
|
||||
| **Reachability policy, network** | open · segmented · asymmetric between segments | yes | `policy:` — inter-segment rules, as a segmented router enforces |
|
||||
| **Segments per gateway** | one · several behind one router | yes | identical gateway declarations mean one gateway machine |
|
||||
| **Gateway state** | permanent · expiring mappings | yes | `mapping_ttl:` |
|
||||
| **Path MTU** | standard · reduced | yes | `segments[].mtu` |
|
||||
| **Overlapping ranges** | distinct · two sites both on `192.168.1.0/24` | yes | segments may carry the same range |
|
||||
| **Segment count** | one · many · isolated island | yes | `kind:` distinguishes an island from a public network |
|
||||
| **Adjacency** | same segment · routed · unrelated public networks | yes | public segments are routed, never bridged — so no two are adjacent unless declared so |
|
||||
| **Gateway depth** | direct · one gateway · nested | **partly** | `to:` chains, so nesting exists; `published:` names one gateway, so forwarding through two does not |
|
||||
| **Address stability** | static · dynamic · changes mid-run | **partly** | a machine can be *moved*; an address changing under it in place cannot be stated |
|
||||
| **Path quality** | latency · loss · bandwidth | **no**, deliberately | changes performance, not correctness — modelling it makes a network simulator, not a fixture |
|
||||
|
||||
### What closing the gaps changed
|
||||
|
||||
**Address family was not a field.** It changed the position model: a machine behind a household
|
||||
gateway is typically *unforwardable on v4 and directly attached on v6, simultaneously*. So
|
||||
reachability is a property of *(machine, family)*, and *"can these two nodes reach each other"*
|
||||
is no longer a yes/no question — it is asked once per family, and the asymmetric answers are the
|
||||
interesting ones. That distinction did not exist in the model an hour ago and would have been
|
||||
discovered by a mesh failing over one family while reporting the other.
|
||||
|
||||
**`inbound:` became necessary because of v6.** With NAT, unreachability was implied by the
|
||||
topology. With a globally routable v6 address, a machine is reachable unless something refuses —
|
||||
so refusing has to be sayable, or v6 addressing would silently imply reachability.
|
||||
|
||||
**`nat:` became a list rather than a boolean** for the same reason: a real gateway translates v4
|
||||
and routes v6, and a boolean cannot say that.
|
||||
|
||||
### What remains open, and whether it matters
|
||||
|
||||
Two partial axes, both extensible when something needs them, neither blocking: **nested
|
||||
forwarding** and **an address changing in place**. A machine can already be moved, which covers
|
||||
the roaming case; what is missing is a lease expiring underneath a machine that stays put.
|
||||
|
||||
One deliberate exclusion: **path quality**. Latency and loss change how fast the mesh is, not
|
||||
whether it is correct. If a timeout turns out to be load-bearing that judgement should be
|
||||
revisited — and it would be revisited by a real failure, which is the right trigger.
|
||||
|
||||
### The honest summary
|
||||
|
||||
The model now covers **where a machine sits**, **what the path between machines is like**, and
|
||||
**what is permitted between them** — across both address families. Together those are what the
|
||||
mesh's reachability logic turns on.
|
||||
|
||||
It contains almost no equipment, by the test above: a device that does not change what an IP
|
||||
packet can do has no fault for a scenario to catch. What it does contain is every device that
|
||||
does — which turns out to be the router, and a modem only when the modem is also a router.
|
||||
|
||||
What it still does not model is *change over time* beyond moving a machine, *degradation* short
|
||||
of failure, and *publishing through two gateways at once*. The first two are chosen; the third
|
||||
is the one real absence, and it is exactly the double-NAT case.
|
||||
|
||||
## Open
|
||||
|
||||
- **`user` and `edge` profiles have no scenario.** A lab machine is always privileged, so the
|
||||
two profiles that exist for unprivileged and phone-like participation cannot currently be
|
||||
exercised. Either the lab grows a way to run the host unprivileged, or those profiles are
|
||||
developed against something that is not a virtual machine.
|
||||
- **Attaching and detaching during a run.** `segment: detached` covers a machine at rest;
|
||||
moving one between segments while a scenario is live is what makes a roaming node
|
||||
interesting, and that is lifecycle rather than declaration.
|
||||
two profiles that exist for unprivileged and phone-like participation cannot be exercised.
|
||||
Either the lab grows a way to run the host unprivileged, or those profiles are developed
|
||||
against something that is not a virtual machine. This is the largest gap.
|
||||
- **Where `place:` gets its artifacts from.** Before the mesh is self-hosting these come from
|
||||
outside; afterwards from the mesh itself. The declaration should not have to care, which
|
||||
suggests a named source rather than a path.
|
||||
- **Multiple scenarios at once.** Each needs its own segments and addresses. Whether the
|
||||
declaration carries absolute addresses, as above, or a template the lab allocates from,
|
||||
decides whether two scenarios can run side by side.
|
||||
- **Nested forwarding** — `published:` names one gateway, so a machine behind two cannot be
|
||||
published through both.
|
||||
- **An address changing in place**, as a DHCP lease expiring under a machine that has not moved.
|
||||
|
||||
@@ -0,0 +1,140 @@
|
||||
---
|
||||
layer: to-be
|
||||
status: designed
|
||||
code: [mesh-lab]
|
||||
updated: 2026-08-23
|
||||
decisions:
|
||||
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
|
||||
- 02-DECISIONS/0031-the-lab-provides-the-underlay.md
|
||||
- 02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md
|
||||
---
|
||||
|
||||
# Scenario lifecycle
|
||||
|
||||
The first thing the lab must do, and the only thing it must do before anything else can be
|
||||
written: **materialise a mesh, return it to a known state, and destroy it**
|
||||
([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)).
|
||||
|
||||
A [declaration](02-scenario-declaration.md) describes a scenario. This describes what happens
|
||||
to one.
|
||||
|
||||
## The verbs
|
||||
|
||||
| Verb | Does |
|
||||
|---|---|
|
||||
| `raise` | materialise a declaration into a running scenario instance |
|
||||
| `snapshot` | name the current state of the whole scenario |
|
||||
| `restore` | return the whole scenario to a named state |
|
||||
| `move` | change a machine's position while the scenario runs |
|
||||
| `exec` | run something on a machine, and get its output |
|
||||
| `destroy` | tear the instance down |
|
||||
|
||||
Six verbs, and `raise` plus `destroy` are enough to be useful. The rest are what make repetition
|
||||
cheap, and cheap repetition is what turns the bootstrap path into an inner development loop
|
||||
rather than a ceremony.
|
||||
|
||||
## Raising, in order
|
||||
|
||||
The order is not arbitrary — each step needs the one before it to exist:
|
||||
|
||||
1. **Segments.** Isolated links, one per declared segment, belonging to this instance and
|
||||
joined to nothing outside it
|
||||
([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)).
|
||||
2. **Gateways.** Derived, never declared as machines: a gateway is materialised for each
|
||||
distinct `gateway:` declaration, sitting on both its segment and its parent, carrying the
|
||||
translation, forwarding and mapping-expiry the declaration asked for.
|
||||
3. **Machines.** Each on its segments, holding its addresses.
|
||||
4. **Policy.** Rules between segments, applied on the gateways that route between them.
|
||||
5. **Placement.** Artifacts onto machines.
|
||||
6. **Snapshot**, if the declaration named one.
|
||||
|
||||
`raise` must wait for the machines to be **usable**, not for the call to return. Measured, those
|
||||
are 14 seconds apart for a single machine — and reporting the earlier number would be the mesh's
|
||||
own recurring failure, transport reported as effect.
|
||||
|
||||
Raising is **convergent, not incremental**: raising an instance that already exists brings it to
|
||||
the declared state rather than failing or duplicating. That is the same model the mesh itself
|
||||
uses, and a lab that behaved differently from the thing it tests would be teaching the wrong
|
||||
habit.
|
||||
|
||||
## A failed raise leaves the wreckage
|
||||
|
||||
A step that fails stops the raise
|
||||
([ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md)) — and **does not tear
|
||||
down**.
|
||||
|
||||
Tearing down on failure destroys the only evidence of what went wrong, which is precisely
|
||||
backwards: a scenario that failed to raise is more interesting than one that succeeded. The
|
||||
instance stays, marked failed, with the step that failed named.
|
||||
|
||||
The caller decides what happens next, and the two callers want different things
|
||||
— the coordinator captures and destroys, a person opens a shell. That is the same
|
||||
one-runner-two-callers split the lab design already makes, applied to failure.
|
||||
|
||||
## Snapshots are whole-scenario
|
||||
|
||||
A snapshot captures **every machine and the state of the network between them**, as one thing.
|
||||
Restoring returns all of it.
|
||||
|
||||
Per-machine snapshots would be cheaper and are wrong. The mesh keeps state that spans nodes —
|
||||
what is assigned where, which grants exist, what has been delivered — so restoring one machine
|
||||
to an earlier moment while its peers move on produces a mesh that has never existed and could
|
||||
not. The faults found there would be artefacts of the lab.
|
||||
|
||||
This is what makes *fresh* and *upgrade* both cheap and both default: one snapshot of a mesh
|
||||
that has never seen a change, another of a mesh running the previous version, and a restore
|
||||
between runs.
|
||||
|
||||
## Moving a machine
|
||||
|
||||
`move` changes a machine's position while the scenario runs: to another segment, with different
|
||||
addresses, or to `detached`.
|
||||
|
||||
It is the roaming case, and it is a **lifecycle** operation rather than a declaration because
|
||||
the interesting part is the transition, not the destination. A mesh that forms correctly with a
|
||||
node at home and correctly with it away may still fail to notice it moved.
|
||||
|
||||
Moving does not invalidate a snapshot. A snapshot is a state to return to; a move is a change
|
||||
made after it, and returning undoes it like any other change.
|
||||
|
||||
## Reaching in
|
||||
|
||||
Everything the lab does to a machine goes through the virtualisation layer, never over IP
|
||||
([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)). `exec` runs a
|
||||
command on a machine and returns its output.
|
||||
|
||||
This has one consequence worth stating plainly: **a reachability question is asked from inside**.
|
||||
*Can this machine reach that one* is `exec` on the first, testing the second — not a probe from
|
||||
the workstation. The workstation is not on the scenario's network and its opinion of
|
||||
reachability would be a different question with a misleadingly similar answer.
|
||||
|
||||
Anything a person wants to open in a browser needs a deliberate forward out of the instance.
|
||||
Nothing leaks by default.
|
||||
|
||||
## The two callers
|
||||
|
||||
The lab design already establishes that the runner serves the coordinator and a person, and
|
||||
that anything only one of them can do will drift. Applied here:
|
||||
|
||||
| | the coordinator | someone working on the mesh |
|
||||
|---|---|---|
|
||||
| wants | structured results, clean teardown, no prompts | readable output, the scenario **left standing**, one thing re-run without repeating the rest |
|
||||
| on failure | capture, then destroy | leave it, open a shell |
|
||||
|
||||
Both use the same verbs. The difference is what happens after the verdict, which is a caller's
|
||||
decision rather than a second implementation.
|
||||
|
||||
## Open
|
||||
|
||||
- **What a snapshot costs — measured, and the answer is a blocker.** On the current host a
|
||||
snapshot is a full copy: 1.6 GB and ten seconds for one small machine, over two minutes when
|
||||
observed again, projecting to roughly a minute and a half for a four-machine reset-and-rerun
|
||||
cycle. The cause is the storage driver, not virtual machines. See
|
||||
[research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md).
|
||||
- **Instance naming.** A declaration is a kind and instances are many; how they are named
|
||||
decides whether a person can find the one they left standing yesterday.
|
||||
- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying
|
||||
the instance must not destroy them.
|
||||
- **Placement before the mesh is self-hosting.** `place:` needs artifacts from somewhere, and
|
||||
before the mesh builds itself that somewhere is outside it — the open question from
|
||||
[research 006](../../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md).
|
||||
@@ -0,0 +1,116 @@
|
||||
---
|
||||
layer: to-be
|
||||
status: designed
|
||||
code: [mesh-lab]
|
||||
updated: 2026-08-24
|
||||
decisions:
|
||||
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
|
||||
- 02-DECISIONS/0008-a-failed-step-fails-the-job.md
|
||||
---
|
||||
|
||||
# Installing the lab on a clean machine
|
||||
|
||||
The lab has prerequisites — a virtualisation daemon, copy-on-write storage, a pool, an identity
|
||||
permitted to talk to it — and it cannot get them from the mesh, because it is where the mesh is
|
||||
built ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)).
|
||||
|
||||
So the lab needs an install path of its own. This describes it, and the shape it has to take is
|
||||
determined by two failures observed while measuring
|
||||
([research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md)).
|
||||
|
||||
## The two failures that shape this
|
||||
|
||||
**One: installed is not available.** The virtualisation package was present and explicitly
|
||||
installed. Both its units were disabled, the operator was in no group, and the client reported
|
||||
the server unreachable. Nothing had failed — the declaration was satisfied exactly as written
|
||||
([`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)).
|
||||
|
||||
**Two, and worse: available is not adequate.** With the storage tooling absent, the daemon
|
||||
offered one driver, everything worked, and snapshots took **seventy-six times longer** than they
|
||||
needed to. Nothing failed. Nothing warned. A lab in that state runs correctly and is simply too
|
||||
slow to use — and the inner loop it exists to provide quietly does not exist.
|
||||
|
||||
The second is the more dangerous shape and it is a variant the mesh has not catalogued before.
|
||||
Its usual failure is *reported success and did nothing*. This is **reported success and did it
|
||||
seventy-six times slower**, which no error surface catches because nothing is wrong.
|
||||
|
||||
## What follows: the lab verifies capability, never installation
|
||||
|
||||
The install path may differ. **The verification does not.**
|
||||
|
||||
Before the lab will raise anything, it asserts the outcomes it needs — not that packages are
|
||||
present, but that the machine can actually do the work:
|
||||
|
||||
| Assertion | Failing means |
|
||||
|---|---|
|
||||
| the daemon answers **as the invoking user**, not as root | a group membership that was granted but never took effect |
|
||||
| a copy-on-write storage driver is offered | the userspace tooling is missing; snapshots will be full copies |
|
||||
| **the pool the lab will use is on that driver** | a pool exists, and is the slow kind — the failure that has no symptom |
|
||||
| hardware virtualisation is present | machines will be emulated and unusably slow |
|
||||
| an image can be fetched or is cached | the first raise will fail late instead of early |
|
||||
|
||||
Each check states **why it matters**, in the terms of what it costs. *"The pool uses the `dir`
|
||||
driver"* means nothing to someone who does not already know it means seventy-six times slower
|
||||
and unbounded at worst.
|
||||
|
||||
**The lab refuses to run degraded.** It does not warn and continue: a warning about a slow inner
|
||||
loop is read once and ignored forever, and the loop stays slow. This is
|
||||
[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) applied where the failure is
|
||||
performance rather than an error.
|
||||
|
||||
## Two ways the prerequisites arrive
|
||||
|
||||
**On a machine the mesh manages** — a module declares them, and a hook turns them into
|
||||
capabilities: the units enabled, the group granted, the pool created on the right driver. This
|
||||
already works; it is what was done for the virtualisation daemon itself.
|
||||
|
||||
**On a machine the mesh does not manage** — the lab's own bootstrap does it. One command, on a
|
||||
clean machine, that installs what is missing and configures it.
|
||||
|
||||
The second path is not a convenience. It is **required**, because the lab must work before the
|
||||
mesh does, and a lab that could only be installed by a mesh would be unable to host the
|
||||
development of the mesh that installs it.
|
||||
|
||||
Both paths end at the same verification. Whoever satisfied the prerequisites, the lab checks
|
||||
them itself — because the lesson of the first failure is precisely that *something else said it
|
||||
was done* is not evidence.
|
||||
|
||||
## The lab is the second thing installed by hand
|
||||
|
||||
Worth stating, because it looks like an exception and is not.
|
||||
|
||||
The node host is the one thing installed by hand on a machine
|
||||
([research 006](../../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)): everything else
|
||||
arrives through it. The lab is the same shape on a workstation — installed once, by hand, and
|
||||
then everything about the mesh is developed inside it.
|
||||
|
||||
Two bootstraps, at two levels, for the same reason: **something has to be first, and pretending
|
||||
otherwise produces a circularity that gets papered over with a script nobody exercises.**
|
||||
|
||||
## What a clean install actually needs
|
||||
|
||||
In order, on a machine with nothing:
|
||||
|
||||
1. **A virtualisation daemon**, running, with its socket enabled.
|
||||
2. **Copy-on-write storage tooling** — the kernel side is usually already present; it is the
|
||||
userspace half that is missing and that decides whether the driver is offered at all.
|
||||
3. **A pool on that driver.** A loop-backed file is sufficient and needs no partitioning, which
|
||||
matters: requiring a dedicated filesystem would make the lab uninstallable on a machine
|
||||
already in use.
|
||||
4. **Group membership** for the operator — which does **not** apply to sessions that already
|
||||
existed. Observed directly: a shell whose process tree predated the grant could not reach
|
||||
the daemon while a fresh lookup showed the membership present. The bootstrap has to say so,
|
||||
or the first thing a person meets is a permission error that looks like a broken install.
|
||||
5. **Verification**, as above, before anything is raised.
|
||||
|
||||
## Open
|
||||
|
||||
- **Whether the lab's bootstrap may install packages at all**, given that the mesh's rules
|
||||
forbid installing by hand. The resolution is probably that the lab's bootstrap *is* the
|
||||
sanctioned mechanism on an unmanaged machine, in the way the mesh's own first-node script is —
|
||||
but that is an argument to record, not to assume.
|
||||
- **What "adequate" means numerically.** The checks above are qualitative. A snapshot-time
|
||||
threshold would catch a copy-on-write pool that is slow for some other reason, and would be a
|
||||
real assertion rather than a proxy.
|
||||
- **Whether the lab should own its pool** rather than using an existing one. Owning it makes the
|
||||
driver guaranteed; sharing it avoids duplicating storage on a machine that already has a pool.
|
||||
@@ -12,6 +12,8 @@ document is written and this one's status becomes `implemented`.
|
||||
| [`00-work-breakdown.md`](00-work-breakdown.md) | How the decomposition gets built, in what order, and where a human must look | [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) |
|
||||
| [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md), [0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) |
|
||||
| [`02-scenario-declaration.md`](02-scenario-declaration.md) | What a scenario declares — the underlay, and what to place on it | [ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md) |
|
||||
| [`03-scenario-lifecycle.md`](03-scenario-lifecycle.md) | What happens to a scenario — raise, snapshot, restore, move, destroy | [ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md) |
|
||||
| [`04-lab-installation.md`](04-lab-installation.md) | Getting the lab onto a clean machine, and why it verifies capability rather than installation | [ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) |
|
||||
|
||||
## Not yet written
|
||||
|
||||
|
||||
Reference in New Issue
Block a user