--- layer: to-be status: designed code: [mesh-lab] updated: 2026-08-23 decisions: - 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md - 02-DECISIONS/0031-the-lab-provides-the-underlay.md - 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md --- # The scenario declaration A scenario is a **declaration of an underlay**, plus what to put on it. It is the interface everything in the lab hangs off, so it is worth getting small. It states what a hosting provider and a home router would provide, and nothing the mesh is responsible for ([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)). ## Three positions a machine can be in The underlay's whole job is to reproduce **where a machine sits relative to the internet**, because that is what the mesh has to cope with and what only production currently exercises. There are three positions, and they are genuinely different: | Position | Reachable from outside | Address | Example | |---|---|---|---| | **Directly attached** | yes, at its own address | fixed, its own | a hosted server | | **Behind a gateway you control** | only through a forwarded port, at the *gateway's* address | private, plus the gateway's public one | a machine at home | | **Behind a gateway you don't control** | **no** | private, and it changes | a laptop on someone else's network | The third is the hard one and the reason this matters. A machine there can dial out and nothing more: it cannot be published, its apparent address belongs to somebody else's router, and that address changes when it moves. Every assumption a mesh makes about reachability breaks there first. A declaration has to be able to say all three, and to move a machine between them. ## The shape ```yaml scenario: roaming-and-published segments: internet: cidr: 203.0.113.0/24 # the simulated public internet, RFC 5737 home: cidr: 192.168.1.0/24 gateway: to: internet address: 203.0.113.50 # what the world sees this network as nat: true elsewhere: # a network we do not control cidr: 198.51.100.0/24 gateway: to: internet address: 203.0.113.80 nat: true machines: anchor: at: { segment: internet, address: 203.0.113.10 } home-server: at: { segment: home, address: 192.168.1.135 } published: - { port: 443, on: home } # DNAT: 203.0.113.50:443 → 192.168.1.135:443 workstation: at: { segment: home, address: 192.168.1.250 } laptop: at: { segment: home, address: 192.168.1.98 } place: all: [host] anchor: [substrate] snapshot: raised ``` ## What each part means, precisely **`segments`** — a broadcast domain with an address range. A segment with no `gateway:` *is* the internet as far as the scenario is concerned. A segment with one sits behind it. **`gateway:`** — how a segment reaches its parent, and this is where the previous version was too thin. It carries three facts, and all three are load-bearing: - `to:` — the parent segment. - `address:` — **the address the outside world sees this network as.** For a household connection this is the public address the ISP hands out. It is not decoration: it is what a peer records as the endpoint when a machine here dials out, and what a public name for a published machine here resolves to. - `nat:` — whether addresses are translated. `true` gives the ordinary household case: many private machines behind one public address. `false` describes a routed range, where machines keep their own addresses and the gateway only forwards. The lab materialises a machine to be the gateway. That is the one implicit machine in an otherwise explicit declaration, and it exists because NAT has to run somewhere. **`machines[].at`** — segment and address. That pair alone determines which of the three positions a machine is in: on a gateway-less segment it is directly attached; on a segment with a gateway it is behind one. **`machines[].published`** — a destination-NAT rule on a named gateway, stated as an outcome rather than a port list. `{ port: 443, on: home }` means the `home` gateway forwards its own `203.0.113.50:443` to this machine's `443`. The resulting public endpoint is derivable, which is the point: a scenario never writes an endpoint down, and the mesh has to discover it. A machine may be published on **any gateway between it and the internet** — which is how *"our LAN also has a public IP"* is expressed, and why `on:` names the gateway rather than being implied. It cannot be published at all on a gateway the scenario models as foreign; attempting it is a declaration error, because that is precisely the constraint being reproduced. **`at: detached`** — on no segment. A machine that exists and can reach nothing. ## Moving a machine is a lifecycle operation `at:` states where a machine *starts*. Moving it is something a run does: ``` move laptop → { segment: elsewhere, address: 198.51.100.23 } move laptop → detached move laptop → { segment: home, address: 192.168.1.98 } ``` This is the roaming case made testable, and it is the one that finds the interesting faults. The same machine, the same identity, three positions in one run: at home where its peers can reach it directly, on a foreign network where it can only dial out and its apparent address belongs to a router it does not control, and asleep. Whether the overlay survives that, re-forms, and is noticed to have changed endpoint is **observed**, never arranged ([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)). ## Why the addresses are load-bearing The internet segment uses RFC 5737 documentation space, and this is not a stylistic choice. The mesh decides *public versus private* by matching the address. A private range on the segment meant to be routable makes a would-be hub test as unreachable, and **the mesh silently never forms** — no error, no failed step, just a mesh that does not exist. Research 004 calls this the single most important fact in its analysis. RFC 5737 reserves three ranges, which is exactly enough for the topology above: | Range | Used for | |---|---| | `203.0.113.0/24` | the internet segment itself — directly attached machines, and gateway addresses | | `198.51.100.0/24` | a foreign network, so a roaming machine's apparent address is plainly not ours | | `192.0.2.0/24` | spare — a second foreign network, or a second site | Private segments use RFC 1918 and can be **byte-identical to production**, because those addresses mean the same thing everywhere. Only the public side is substituted, and only because it must be. The format should make getting this wrong hard rather than merely documented: a segment without a `gateway:` is a public segment, and an address in it — including a gateway's `address:` — that is not documentation space is a declaration error, refused before anything is raised. That is [ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) applied to a configuration file: the failure it prevents is silent, so the check has to be loud. ## The same declaration serves both classes The bootstrap and full scenarios differ **only in `place:`** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)). Everything about the underlay is identical, which is what makes one a strict subset of the other rather than a fork. ```yaml # bootstrap — tiers 0 and 1 place: all: [host] anchor: [substrate] # full — adds a control plane, a forge, and a module under test place: all: [host] anchor: [substrate, control, forge] module: a-web-service assert: - the service answers on its published name - the certificate presented is valid for that name ``` `module:` and `assert:` are meaningless in a bootstrap scenario and absent from one. A bootstrap scenario's verdict comes from what the host reports about the state it reconciled, not from an assertion runner — which is why assertion execution is second in the build order, not first. ## What a scenario deliberately cannot say - **Overlay addresses, the hub, peer configuration.** Outcomes, not inputs ([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)). - **What a machine is in mesh terms** — server or workstation, its site, its names. Mesh configuration, established by the mesh. - **A host's capability profile.** Detected, never declared. - **Steps.** A scenario is a desired state. Anything expressed as an ordered list of actions belongs in the lifecycle, not the declaration. ## Open - **`user` and `edge` profiles have no scenario.** A lab machine is always privileged, so the two profiles that exist for unprivileged and phone-like participation cannot be exercised. Either the lab grows a way to run the host unprivileged, or those profiles are developed against something that is not a virtual machine. This is the largest gap. - **Where `place:` gets its artifacts from.** Before the mesh is self-hosting these come from outside; afterwards from the mesh itself. The declaration should not have to care, which suggests a named source rather than a path. - **Multiple scenarios at once.** Each needs its own segments and addresses, and the shape above writes addresses absolutely. Whether a scenario carries literal addresses or a template the lab allocates from decides whether two can run side by side — and there are only three documentation ranges to go round. - **Gateway behaviour beyond forwarding.** A real household gateway also has a NAT table with timeouts, and connection tracking that drops idle flows. Whether a scenario can express *"the gateway forgets a mapping after N seconds"* decides whether keepalive behaviour is testable or merely hoped for.