ADR 0031 and the scenario declaration

The lab provides the underlay; the mesh builds the overlay. This is the
boundary that decides whether the lab is worth having: a scenario that
assigns overlay addresses, elects the hub and writes peer configuration
certifies its own work — if the mesh's peering is broken, that scenario
still comes up green. The most valuable thing the lab can test is exactly
the part pre-building would replace.

So a scenario declares what a hosting provider and a home router would
provide: segments, which machine sits where at which address, what NAT is
between them, which ports are forwarded, which machines are detached. It
declares nothing about overlay addresses, hubs, peering, names or
certificates, all of which become outcomes to observe.

The declaration has four parts — segments, machines, place, snapshot — and
the two scenario classes differ only in place. That is what makes one a
strict subset of the other rather than a fork.

Research 004's most important finding becomes a format constraint rather
than a footnote: the routable segment must use RFC 5737 documentation
space, because the mesh decides public versus private by matching the
address, and a private range there makes the hub test as unreachable while
the mesh silently never forms. A segment without behind: is routable, and a
non-documentation address in it should be refused before anything is
raised — ADR 0008 applied to a configuration file, since the failure it
prevents has no error at all.

Four things left open, including the one that matters most: a lab machine
is always privileged, so the user and edge profiles have no scenario that
exercises them.
This commit is contained in:
2026-08-23 22:33:52 +02:00
parent 4d387d2998
commit a72fea5342
4 changed files with 231 additions and 2 deletions
+1 -1
View File
@@ -2,7 +2,7 @@
status: active
initiated: 2026-08-22
touches: [03-DESIGN/00-as-is/01-mesh-and-transport.md, 03-DESIGN/01-to-be/01-end-to-end-testing.md]
became: [02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md]
became: [02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md, 02-DECISIONS/0031-the-lab-provides-the-underlay.md, 03-DESIGN/01-to-be/02-scenario-declaration.md]
---
# 004 — Reproducing the mesh network in a lab
@@ -0,0 +1,87 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 31. The lab provides the underlay; the mesh builds the overlay
## Context
[Research 004](../01-RESEARCH/004-lab-network/analysis.md) worked out the topology a lab has to
reproduce: a routable segment using documentation addresses, a household segment behind NAT, a
router that forwards exactly one port so a *published-but-NATed* node is real, and a machine
that can attach to either segment or detach entirely.
It also records what makes that topology **mean** something, and this is where a boundary has
to be drawn. Hub election is by convention rather than by flag — the hub is the node whose
profile is server and whose overlay address begins `10.10.0.1`. Direct peering depends on two
nodes sharing a site. Names resolve from mesh configuration on each node.
Those are all facts the *mesh* establishes. The question is whether a scenario declares them.
It is tempting to say yes, because a scenario that hands you a working overlay is a scenario
you can start testing against immediately.
## Considered options
1. **The lab configures the overlay too** — assign the overlay addresses, elect the hub, write
the peer configuration, seed the names. Rejected, and the reason is the whole point of the
lab: **a lab that builds the overlay certifies its own work.** If the mesh's peering logic
is broken, a scenario that pre-built the peering still comes up green. The most valuable
thing the lab can test is precisely the part this would replace.
2. **The lab provides nothing but bare machines** — no addressing, no segments, no NAT. Also
rejected. Then the scenario cannot reproduce *published but behind NAT*, which research 004
identifies as the case that only exists in production today, and the lab loses its reason to
use virtual machines at all.
3. **The lab provides the underlay; the mesh builds the overlay.** Chosen.
## Decision
**A scenario declares the underlay** — the facts a machine would have before any of our
software touched it:
- which segments exist, and their address ranges
- which machine sits on which segment, at which address
- what NAT sits between them, and which ports are forwarded through it
- which machines are detached, and can be attached or detached during a run
**A scenario declares nothing about the overlay** — no overlay addresses, no hub, no peering,
no names, no certificates. Those are the mesh's job, and a scenario that supplied them would be
testing itself.
The rule stated in one line: **a scenario provides what a hosting provider and a home router
would provide, and nothing our software is responsible for.**
## Consequences
- **The overlay becomes a thing under test rather than a fixture.** Whether peers form,
whether the hub is elected, whether a NATed node's endpoint is learned — all of it is
observed rather than arranged. That is the class of fault research 004 says is discoverable
only in production today.
- The lab stays small, and stays honest. It needs to know about virtualisation, bridges,
addresses and NAT. It never needs to know what a mesh node is.
- A scenario cannot assert "the overlay came up" as a precondition, because it is an outcome.
A bootstrap scenario that wants a working overlay has to wait for one and check, which is
the correct shape.
- **The address ranges are load-bearing, not cosmetic.** The routable segment uses RFC 5737
documentation space specifically because the mesh's own code decides *public versus private*
by matching the address — a private range there makes the hub test as unreachable, and the
mesh silently never forms. Research 004 calls this the single most important fact in the
document, and the declaration format has to make getting it wrong hard.
- The router is a machine the lab materialises without being asked, because NAT requires
somewhere to run. That is an implicit machine in an otherwise explicit declaration, and it is
worth knowing about rather than discovering.
- **Host capability profiles are detected, not declared** — a consequence of
[research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md), and consistent here: a
scenario does not say what a machine is allowed to do, it provides a machine. Which leaves an
open question: a lab machine is always privileged, so the `user` and `edge` profiles have no
scenario that exercises them yet.
## References
- [Research 004](../01-RESEARCH/004-lab-network/analysis.md) — the topology, the documentation
ranges, and the hub-election and peering conventions this deliberately does not touch.
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the two scenario classes this
declaration has to serve without forking.
@@ -0,0 +1,141 @@
---
layer: to-be
status: designed
code: [mesh-lab]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
- 02-DECISIONS/0031-the-lab-provides-the-underlay.md
- 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
---
# The scenario declaration
A scenario is a **declaration of an underlay**, plus what to put on it. It is the interface
everything in the lab hangs off, so it is worth getting small.
It states what a hosting provider and a home router would provide, and nothing the mesh is
responsible for ([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)).
## The shape
```yaml
scenario: published-behind-nat
segments:
wan:
cidr: 203.0.113.0/24 # RFC 5737 — never routes on the real internet
lan:
cidr: 192.168.1.0/24
behind: wan # NAT; the lab materialises a router
machines:
anchor:
segment: wan
address: 203.0.113.10
home-server:
segment: lan
address: 192.168.1.135
forwarded: [443] # reachable from wan through the router
workstation:
segment: lan
address: 192.168.1.250
laptop:
segment: detached # reachable by nothing until attached
place:
all: [host]
anchor: [substrate]
snapshot: raised
```
That is a complete bootstrap scenario. Nothing in it mentions the overlay, a hub, peering,
names or certificates — all of which are outcomes to be observed.
## The four parts
**`segments`** — the networks that exist. `behind:` declares NAT, and is the only place a
router comes from: the lab materialises one without being asked, because NAT has to run
somewhere. This is the one implicit machine in an otherwise explicit declaration.
**`machines`** — what sits where. A machine has a segment and an address, and that is nearly
all. `forwarded:` opens a port through the router, which is what makes *published but behind
NAT* reproducible — the case that exists only in production today. `segment: detached` is a
machine on no network, which is how a roaming node is expressed at rest.
**`place`** — what goes inside. `all:` applies to every machine; a machine name overrides for
that machine. This is the only part that differs between the two scenario classes.
**`snapshot`** — names the state once placement finishes, so a run can return to it without
raising everything again. Snapshots are what make repetition cheap, and cheap repetition is
what makes the bootstrap path the inner development loop rather than a ceremony.
## Why the addresses are load-bearing
The routable segment uses RFC 5737 documentation space, and this is not a stylistic choice.
The mesh decides *public versus private* by matching the address. A private range on the
segment meant to be routable makes the hub test as unreachable, and **the mesh silently never
forms** — no error, no failed step, just a mesh that does not exist. Research 004 calls this
the single most important fact in its analysis.
So the format should make this hard to get wrong rather than merely documented: a segment
without `behind:` is a routable segment, and an address in it that is not documentation space
is a declaration error, refused before anything is raised. That is
[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) applied to a configuration
file — the failure it prevents is silent, so the check has to be loud.
## The same declaration serves both classes
The bootstrap and full scenarios differ **only in `place:`**
([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)). Everything
about the underlay is identical, which is what makes one a strict subset of the other rather
than a fork.
```yaml
# bootstrap — tiers 0 and 1
place:
all: [host]
anchor: [substrate]
# full — adds a control plane, a forge, and a module under test
place:
all: [host]
anchor: [substrate, control, forge]
module: a-web-service
assert:
- the service answers on its published name
- the certificate presented is valid for that name
```
`module:` and `assert:` are meaningless in a bootstrap scenario and absent from one. A
bootstrap scenario's verdict comes from what the host reports about the state it reconciled,
not from an assertion runner — which is why assertion execution is second in the build order,
not first.
## What a scenario deliberately cannot say
- **Overlay addresses, the hub, peer configuration.** Outcomes, not inputs
([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)).
- **What a machine is in mesh terms** — server or workstation, its site, its names. Mesh
configuration, established by the mesh.
- **A host's capability profile.** Detected, never declared.
- **Steps.** A scenario is a desired state. Anything expressed as an ordered list of actions
belongs in the lifecycle, not the declaration.
## Open
- **`user` and `edge` profiles have no scenario.** A lab machine is always privileged, so the
two profiles that exist for unprivileged and phone-like participation cannot currently be
exercised. Either the lab grows a way to run the host unprivileged, or those profiles are
developed against something that is not a virtual machine.
- **Attaching and detaching during a run.** `segment: detached` covers a machine at rest;
moving one between segments while a scenario is live is what makes a roaming node
interesting, and that is lifecycle rather than declaration.
- **Where `place:` gets its artifacts from.** Before the mesh is self-hosting these come from
outside; afterwards from the mesh itself. The declaration should not have to care, which
suggests a named source rather than a path.
- **Multiple scenarios at once.** Each needs its own segments and addresses. Whether the
declaration carries absolute addresses, as above, or a template the lab allocates from,
decides whether two scenarios can run side by side.
+2 -1
View File
@@ -10,7 +10,8 @@ document is written and this one's status becomes `implemented`.
| Document | Covers | Rests on |
|---|---|---|
| [`00-work-breakdown.md`](00-work-breakdown.md) | How the decomposition gets built, in what order, and where a human must look | [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) |
| [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md) |
| [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md), [0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) |
| [`02-scenario-declaration.md`](02-scenario-declaration.md) | What a scenario declares — the underlay, and what to place on it | [ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md) |
## Not yet written