Files
hq/03-DESIGN/01-to-be/02-scenario-declaration.md
T
jschoubben 573a94e102 Correct the record: a limitation that no longer exists, and one that was never written
The connectivity design still said a hub cannot be filtered — a gap recorded in
the morning and closed in the afternoon, left standing as though it were
current. Worse than a stale date: it would send somebody away from something
that works.

`restart-on` was described nowhere, including the part added today that lets a
service reflect a file another module put on the machine. A rule the host
enforces and no document mentions is a rule nobody can rely on.

And nine of fifteen design documents claimed an `updated:` older than their last
change, some by a week. That field is what cross-cutting views are generated
from, so it is not decoration.
2026-08-31 12:33:35 +02:00

645 lines
32 KiB
Markdown

---
layer: to-be
status: in-progress
code: [mesh-lab]
updated: 2026-08-28
decisions:
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
---
# The scenario declaration
A scenario is a **declaration of an underlay**, plus what to put on it. It is the interface
everything in the lab hangs off, so it is worth getting small.
It states what a hosting provider and a home router would provide, and nothing the mesh is
responsible for ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
## Public networks are unrelated, and routed rather than bridged
The internet is not a network. It is a very large number of unrelated networks that route to
each other, and a machine in one is **many hops** from a machine in another with no shared
broadcast domain between them.
So a scenario does not have *an* internet segment. It has **one public segment per public
network**, each with its own unrelated prefix, and the lab wires them together **through a
router, never onto a shared bridge**.
That distinction is load-bearing, and putting several public addresses in one prefix would
quietly make four things true that are false in reality:
| If public addresses share a segment | Reality |
|---|---|
| machines resolve each other by ARP and talk directly | they are routed, many hops apart |
| TTL never decrements | every hop decrements it |
| broadcast and multicast reach across | neither crosses a router |
| any two are adjacent | adjacency is the exception, not the rule |
The third is not hypothetical here. The mesh has already been bitten by multicast name
resolution — [`00-as-is/01-mesh-and-transport.md`](../00-as-is/01-mesh-and-transport.md)
records that mesh names are deliberately not multicast names, after a delay and a
one-node-only failure mode. A lab where "the internet" is one broadcast domain would let a node
discover a peer by multicast that it could never discover in production, and report success.
**The rule: `kind: public` segments are routed to one another and never bridged.** It is a
property of how the lab wires them, not a field anyone sets, because there is no correct
scenario in which two public networks are adjacent.
### Which addresses to use
RFC 5737 reserves three ranges and RFC 3849 reserves one v6 prefix. Each public network takes
its own, and they are chosen to look nothing like each other — because in reality they would
not:
| Public network | v4 | v6 |
|---|---|---|
| a hosting provider | `192.0.2.0/24` | `2001:db8:a::/48` |
| a household ISP | `198.51.100.0/24` | `2001:db8:b::/48` |
| a mobile or foreign network | `203.0.113.0/24` | `2001:db8:c::/48` |
Three is enough for the topologies that matter, and a fourth public network can subnet one of
them — an ISP handing out `198.51.100.0/25` and `198.51.100.128/25` to two customers is exactly
what really happens.
Private segments still use RFC 1918 and stay byte-identical to production.
## What NAT does, and why the design turns on it
A household or office has **one** address the outside world can see, and **many** machines
behind it. Network address translation is what reconciles those.
When a machine inside dials out, the gateway rewrites the packet's source from the private
address to the public one, **remembers the mapping**, and rewrites the replies on the way back.
Four consequences follow, and every one of them shapes this design:
1. **Outbound works; inbound does not.** A mapping exists only because something inside started
a conversation. Nothing outside can start one — there is no mapping to look up, and no way
to know which internal machine was meant.
2. **A forwarded port is a permanent mapping made by hand**, in the inbound direction:
*anything arriving at the public address on 443 goes to this machine.* That is the only way
a machine behind NAT becomes reachable, and it requires control of the gateway.
3. **Mappings expire.** A gateway forgets one that goes unused. This is why anything holding a
connection through NAT sends keepalives, and why a mesh that does not is fine until it is
idle.
4. **From outside, every machine behind the gateway looks like one address.** Identity and
address stop corresponding.
This is why the mesh dials outward and never inward
([ADR 0002](../../02-DECISIONS/0002-nodes-communicate-over-a-broker.md)), why a hub exists at
all, and why a node's endpoint is something a peer **learns** from arriving packets rather than
something anyone configures.
**Carrier-grade NAT** is the same mechanism applied by an ISP: your own gateway gets a private
address too, and the public one is shared with strangers. Nothing can be forwarded, because the
rule would have to live on equipment you do not own. Common on mobile connections and
increasingly on fixed ones.
## Three positions a machine can be in
The underlay's whole job is to reproduce **where a machine sits relative to the internet**,
because that is what the mesh has to cope with and what only production currently exercises.
There are three positions, and they are genuinely different:
| Position | Reachable from outside | Apparent address | Example |
|---|---|---|---|
| **Attached** | yes, at its own address | its own | a hosted server |
| **Behind a forwardable gateway** | only through a forwarded port, at the *gateway's* address | the gateway's | a machine at home |
| **Behind an unforwardable gateway** | **no** | someone else's, and it changes | a laptop on a café network; anything behind carrier-grade NAT |
The axis is **forwardability, not ownership** — which is worth stating because the obvious
framing gets it wrong. Carrier-grade NAT is *your* connection and is still unforwardable, so it
belongs in the third row alongside the café. What the mesh has to cope with is whether an
inbound mapping can be made, not who owns the equipment.
The third position is the hard one. A machine there can dial out and nothing more: it cannot be
published, its apparent address belongs to a router it does not control, and that address
changes when it moves. Every assumption a mesh makes about reachability breaks there first.
A declaration has to be able to say all three, and to move a machine between them.
## Reachability is per address family, not per machine
Adding IPv6 is not a field. It changes the position model, and the reason is worth stating
before the syntax.
**IPv6 usually has no NAT.** A machine behind a household gateway can hold a *globally routable*
v6 address while its v4 address is private and unforwardable. The same machine, at the same
moment, is in **two different positions at once**:
| | IPv4 | IPv6 |
|---|---|---|
| a typical machine at home | behind an unforwardable-or-forwardable gateway | **attached**, directly reachable |
| a machine on mobile data | behind carrier-grade NAT | often attached, sometimes absent entirely |
| a machine on an older network | attached or behind NAT | **no address at all** |
So the three positions apply **per family**, and a machine's reachability is a property of
*(machine, family)* rather than of the machine. A mesh that treats reachability as one fact per
node will reach a peer over one family, fail over the other, and report whichever it tried.
That has a direct consequence for what the lab is for: *"can these two nodes reach each other"*
stops being a yes/no question. It is asked once per family, and the interesting answers are the
asymmetric ones.
The v6 documentation prefix is `2001:db8::/32` (RFC 3849) — the exact counterpart of the RFC
5737 rule, and load-bearing for the same reason.
## The shape
```yaml
scenario: roaming-and-published
segments:
hosting: # one public network
kind: public
cidr: [192.0.2.0/24, 2001:db8:a::/48]
isp-home: # another, unrelated — routed to it, not bridged
kind: public
cidr: [198.51.100.0/24, 2001:db8:b::/48]
home:
kind: private
cidr: [192.168.1.0/24, 2001:db8:b:1::/64]
mtu: 1500
gateway:
to: isp-home
address: [198.51.100.7, 2001:db8:b::7] # what the world sees this network as
nat: [v4] # v4 is translated; v6 is routed, not translated
forwardable: true
mapping_ttl: 120s # an unused inbound mapping is forgotten after this
cafe: # a network we do not control
kind: private
cidr: [10.50.0.0/16] # v4 only — no v6 offered here at all
mtu: 1400 # a tunnelled path, smaller than standard
gateway:
to: hosting # it reaches the world via a different public network
address: [192.0.2.200]
nat: [v4]
forwardable: false # café wifi, or carrier-grade NAT
mapping_ttl: 30s # aggressive, as carrier NAT tends to be
machines:
anchor:
at: { segment: hosting, address: [192.0.2.10, 2001:db8:a::10] }
home-server:
at: { segment: home, address: [192.168.1.135, 2001:db8:b:1::135] }
published:
- { port: 443, on: home } # v4 only: 198.51.100.7:443 → 192.168.1.135:443
inbound: allow # v6 is routable here, so this decides whether it is reachable
workstation:
at: { segment: home, address: [192.168.1.250, 2001:db8:b:1::250] }
inbound: deny # a host firewall: dials out, accepts nothing
laptop:
at: { segment: home, address: [192.168.1.98, 2001:db8:b:1::98] }
border: # a machine on two segments at once
at:
- { segment: home, address: [192.168.1.2] }
- { segment: isp-home, address: [198.51.100.60] }
place:
all: [host]
anchor: [substrate]
snapshot: raised
```
## What each part means, precisely
**`segments`** — a broadcast domain with an address range, and a `kind:`. A segment is a single
broadcast domain, which is precisely why several public networks cannot be one segment.
`kind: public` marks a segment that stands in for a public network — and there is normally more
than one, unrelated to each other. `kind: private` is everything else. This is stated rather than inferred, and the earlier version inferred it — *a segment with
no gateway is the internet* — which made an **isolated network inexpressible**: a LAN with no
route out is a private segment with no gateway, and would have been read as public and forced
to use documentation addresses. A mesh spanning a site with no internet access is a real
topology, and the model has to be able to say it.
**`gateway:`** — how a segment reaches its parent, and this is where the previous version was
too thin. It carries three facts, and all three are load-bearing:
- `to:` — the parent segment.
- `address:` — **the address the outside world sees this network as.** For a household
connection this is the public address the ISP hands out. It is not decoration: it is what a
peer records as the endpoint when a machine here dials out, and what a public name for a
published machine here resolves to.
- `nat:` — **which families are translated**, as a list. `[v4]` is the ordinary modern case:
v4 translated, v6 routed. `[v4, v6]` describes a gateway that translates both, which exists
and is worth being able to reproduce. `[]` is a routed range, where machines keep their own
addresses and the gateway only forwards.
- `forwardable:` — whether an inbound mapping can be created. Independent of `nat:`, and the
field that separates a home gateway from carrier-grade NAT. Publishing through a gateway with
`forwardable: false` is a declaration error, because that is exactly the constraint being
reproduced.
- `mapping_ttl:` — how long an unused inbound mapping survives. This is what makes keepalive
behaviour testable: a mesh that holds a connection through NAT without refreshing it works
perfectly until the far side goes quiet for longer than this. Aggressive values reproduce
carrier NAT; omitting it means mappings never expire, which no real gateway does.
**`segments[].mtu`** — the largest packet the segment carries, defaulting to 1500. Lower values
reproduce tunnelled and PPPoE paths. This matters because an overlay adds its own header: a
tunnel over a 1400-byte path establishes a connection and then silently drops large packets,
which is the shape of fault this whole effort exists to stop shipping.
**`machines[].inbound`** — `allow` or `deny`, a host firewall. Distinct from NAT and behaves
differently: a machine can be perfectly routable and still refuse everything unsolicited, which
is the normal state of a v6-addressed machine. Without this, v6 addressing would imply
reachability, and it does not.
The lab materialises a machine to be the gateway. That is the one implicit machine in an
otherwise explicit declaration, and it exists because NAT has to run somewhere.
It is a **container, not a virtual machine** — a router is scenery rather than something under
test, so the fidelity argument that makes a node a virtual machine does not reach it
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). What a router must
reproduce is kernel behaviour, and a container has the same kernel.
**`machines[].at`** — segment and addresses, or a **list** of them for a machine on several
segments at once. Multi-homing is not exotic: it is what a border machine is, and what any node
with both a LAN and a WAN interface is. Each entry carries the addresses that machine holds on
that segment, one per family.
Position follows from the pair, per family: on a `kind: public` segment a machine is attached;
on a private one it is behind that segment's gateway, unless the gateway does not translate
that family — in which case it is attached on that family and behind a gateway on the other.
**`machines[].published`** — a destination-NAT rule on a named gateway, stated as an outcome
rather than a port list. `{ port: 443, on: home }` means the `home` gateway forwards its own
`203.0.113.50:443` to this machine's `443`. The resulting public endpoint is derivable, which is
the point: a scenario never writes an endpoint down, and the mesh has to discover it.
A machine may be published on **any gateway between it and a public network** — which is how *"our
LAN also has a public IP"* is expressed, and why `on:` names the gateway rather than being
implied. It cannot be published at all on a gateway the scenario models as foreign; attempting
it is a declaration error, because that is precisely the constraint being reproduced.
**`at: detached`** — on no segment. A machine that exists and can reach nothing.
## Moving a machine is a lifecycle operation
`at:` states where a machine *starts*. Moving it is something a run does:
```
move laptop → { segment: elsewhere, address: 198.51.100.23 }
move laptop → detached
move laptop → { segment: home, address: 192.168.1.98 }
```
This is the roaming case made testable, and it is the one that finds the interesting faults.
The same machine, the same identity, three positions in one run: at home where its peers can
reach it directly, on a foreign network where it can only dial out and its apparent address
belongs to a router it does not control, and asleep.
Whether the overlay survives that, re-forms, and is noticed to have changed endpoint is
**observed**, never arranged
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
## Why the addresses are load-bearing
The internet segment uses RFC 5737 documentation space, and this is not a stylistic choice.
The mesh decides *public versus private* by matching the address. A private range on the
segment meant to be routable makes a would-be hub test as unreachable, and **the mesh silently
never forms** — no error, no failed step, just a mesh that does not exist. Research 004 calls
this the single most important fact in its analysis.
Which range goes where is covered above, under *public networks are unrelated*: one range per
public network, chosen to look nothing like each other.
Private segments use RFC 1918 and can be **byte-identical to production**, because those
addresses mean the same thing everywhere. Only the public side is substituted, and only because
it must be.
The format should make getting this wrong hard rather than merely documented: a segment without
a `gateway:` is a public segment, and an address in it — including a gateway's `address:` — that
is not documentation space is a declaration error, refused before anything is raised. That is
[ADR 0010](../../02-DECISIONS/0010-delivery.md) applied to a configuration
file: the failure it prevents is silent, so the check has to be loud.
## The same declaration serves both classes
The bootstrap and full scenarios differ **only in `place:`**
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). Everything
about the underlay is identical, which is what makes one a strict subset of the other rather
than a fork.
```yaml
# bootstrap — tiers 0 and 1
place:
all: [host]
anchor: [substrate]
# full — adds a control plane, a forge, and a module under test
place:
all: [host]
anchor: [substrate, control, forge]
module: a-web-service
assert:
- the service answers on its published name
- the certificate presented is valid for that name
```
`module:` and `assert:` are meaningless in a bootstrap scenario and absent from one. A
bootstrap scenario's verdict comes from what the host reports about the state it reconciled,
not from an assertion runner — which is why assertion execution is second in the build order,
not first.
## What a scenario deliberately cannot say
- **Overlay addresses, the hub, peer configuration.** Outcomes, not inputs
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
- **What a machine is in mesh terms** — server or workstation, its site, its names. Mesh
configuration, established by the mesh.
- **A host's capability profile.** Detected, never declared.
- **Steps.** A scenario is a desired state. Anything expressed as an ordered list of actions
belongs in the lifecycle, not the declaration.
## What the model deliberately does not contain
A real network is full of equipment: a modem, a router, switches, access points, controllers.
Almost none of it appears here, and the omission is deliberate rather than an oversight.
**The test is whether a device changes what an IP packet can do.** If two machines can exchange
packets, at the same addresses, with the same reachability and the same MTU, whether or not the
device exists — then the device is invisible to the mesh, and modelling it would add a fixture
with no fault to catch.
| Equipment | Modelled? | Why |
|---|---|---|
| **Switch** | no | Moves frames within a segment. Two machines on a switch are two machines on a segment. |
| **Access point** | no | Bridges wireless clients onto a segment. A machine on wifi and a machine on cable are the same machine to IP. |
| **Network controller** | no | Configures equipment. Its effects appear as segments and policy; it has no packets of its own. |
| **Cabling, PoE, uplink speed** | no | Change performance and availability, not reachability. |
| **Router / security gateway** | **yes — it *is* the gateway** | Translation, forwarding and inter-segment policy all live here. |
| **ISP modem** | **only in router mode** | In bridge mode it is a media converter and invisible. Doing its own NAT, it is a second gateway — and that is double NAT. |
| **VLANs** | **yes — they are segments** | Machines on separate VLANs cannot reach each other except through the router, which is the definition of a separate segment. |
| **Inter-VLAN firewall rules** | **yes** — see below | A rule stopping one segment reaching another is a reachability fact, and the mesh will hit it. |
The two entries worth dwelling on are the ones where a common household setup produces
something the mesh has to survive.
**A modem in router mode gives you double NAT.** Your gateway holds a private address from the
modem, which holds the public one. Forwarding then requires a rule on *both*, and one of them
may not be configurable. This is expressible as nested segments — `to:` chains — but publishing
through two gateways is not, and remains open.
**Segmented networks are the common case, not the exotic one.** A router with separate networks
for trusted machines, guests and devices is ordinary, and the rules between them are ordinary
too. A mesh node on one segment and a mesh node on another are, as far as reachability goes, on
different networks that happen to share a gateway.
## Segments may share a gateway
Several segments can name the same parent and the same external address. That is one router
with several networks behind it, which is what a VLAN-capable gateway is:
```yaml
segments:
trusted:
kind: private
cidr: [192.168.1.0/24]
gateway: { to: isp-home, address: [198.51.100.7], nat: [v4], forwardable: true }
devices:
kind: private
cidr: [192.168.30.0/24]
gateway: { to: isp-home, address: [198.51.100.7], nat: [v4], forwardable: true }
```
Identical gateway declarations mean **one gateway machine**, not two. The lab materialises a
single router serving both, because that is what the topology being reproduced is — and two
routers sharing one address would not work anyway.
## Policy between segments
Sharing a gateway does not mean segments can reach each other. What they may do is stated
separately, because it is a fact about a pair rather than a property of either:
```yaml
policy:
- { from: devices, to: trusted, allow: false } # devices may not initiate to trusted
- { from: trusted, to: devices, allow: true } # the reverse is fine
```
Default is `true` between segments behind the same gateway, matching a router with no rules
configured. Asymmetry is the normal case and the reason this cannot be a single flag: the
useful configuration is almost always one-directional.
This is **not** the same as `machines[].inbound`, and conflating them loses a real distinction:
| | Enforced by | Blocks |
|---|---|---|
| `policy` | the gateway, between segments | everything crossing, regardless of what the destination thinks |
| `inbound` | the machine itself | unsolicited traffic that already reached it |
A mesh node behind a `policy` deny cannot be reached even by a peer that knows exactly where it
is — and that is a real topology, not a contrived one.
## Worked example — a whole mesh of the ordinary kind
The shape research 004 identified: one machine with a routable address, one publicly named but
behind a household connection, one stationary machine on that network, one that roams. Written
out completely, with every field the model has.
```yaml
scenario: the-ordinary-shape
segments:
# ---- three unrelated public networks. Routed to each other, never bridged. ----
hosting: # where the always-on machine lives
kind: public
cidr: [192.0.2.0/24, 2001:db8:a::/48]
isp-home: # the household's uplink
kind: public
cidr: [198.51.100.0/24, 2001:db8:b::/48]
isp-mobile: # wherever the roaming machine happens to be
kind: public
cidr: [203.0.113.0/24, 2001:db8:c::/48]
# ---- private networks behind them ----
home: # the household network
kind: private
cidr: [192.168.1.0/24, 2001:db8:b:1::/64]
mtu: 1492 # PPPoE on the uplink; 1500 if the line is not PPPoE
gateway:
to: isp-home
address: [198.51.100.7, 2001:db8:b::7]
nat: [v4] # v4 translated, v6 routed — the modern default
forwardable: true # the household router is ours to configure
mapping_ttl: 120s
devices: # optional: a segmented network on the same router
kind: private
cidr: [192.168.30.0/24]
gateway:
to: isp-home
address: [198.51.100.7, 2001:db8:b::7] # identical → the SAME gateway machine
nat: [v4]
forwardable: true
mapping_ttl: 120s
cafe: # a network we do not control
kind: private
cidr: [10.50.0.0/16] # RFC 1918 — someone else's private range
mtu: 1400
gateway:
to: isp-mobile
address: [203.0.113.129]
nat: [v4]
forwardable: false # carrier-grade NAT, or simply not ours
mapping_ttl: 30s
policy:
- { from: devices, to: home, allow: false }
- { from: home, to: devices, allow: true }
machines:
anchor: # routable, nothing in front of it
at: { segment: hosting, address: [192.0.2.10, 2001:db8:a::10] }
inbound: allow
home-server: # publicly named, behind the household connection
at: { segment: home, address: [192.168.1.135, 2001:db8:b:1::135] }
published:
- { port: 443, on: home } # v4 reaches it only through the forward
inbound: allow # and v6 reaches it directly, so this matters
workstation: # on the household network, not published
at: { segment: home, address: [192.168.1.250, 2001:db8:b:1::250] }
inbound: deny
laptop: # starts at home; moves during the run
at: { segment: home, address: [192.168.1.98, 2001:db8:b:1::98] }
inbound: deny
place:
all: [host]
anchor: [substrate]
snapshot: raised
```
### What the ISP modem is doing here
**Nothing, if it is in bridge mode** — which is the ordinary arrangement when the household has
its own router. A bridged modem is a media converter: it changes the physical medium and leaves
the packets alone, so it creates no IP-level fact and appears nowhere above.
Were it in router mode it would be a second gateway, `home` would sit behind it rather than
behind `internet` directly, and publishing would need a rule on both — the case the model
cannot yet express.
### What is deliberately absent
Nothing here mentions overlay addresses, which node is the hub, who peers with whom, any name,
or any certificate. Research 004 recorded all of those for this topology, and **a scenario must
not state them** ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)): they
are what the mesh does, and a scenario that supplied them would be certifying its own work.
The absence is the point. Given the declaration above, whether a hub is elected, whether the
NATed machine's endpoint is learned, whether the roaming machine re-forms after moving — all of
it is observed.
### Running it
```
move laptop → { segment: cafe, address: [10.50.3.23] } # it leaves the house
move laptop → detached # it sleeps
move laptop → { segment: home, address: [192.168.1.98, 2001:db8:b:1::98] }
```
One identity, four positions, one run. The `mapping_ttl: 30s` on `cafe` means a connection held
without refreshing dies while it is out — which is the point of putting a number there.
Note that the laptop's three positions are on **three different public networks**: at home it
appears as `198.51.100.7`, at the café as `203.0.113.129`, and the machine it is trying to reach
is on a fourth. None of them are adjacent, all of them are routed. That is the situation the
mesh actually faces.
## Is this general? — the axes a setup can vary along
The question that matters is not *does this cover our mesh*, but **can it express any mesh**.
The standard applied is not "every property a network has". It is **every property that changes
how the mesh behaves**. Bandwidth does not change correctness; MTU does.
| Axis | Values | Expressible | |
|---|---|---|---|
| **Reachability** | attached · forwardable gateway · unforwardable gateway · isolated | yes | the core of the model, and **per family** |
| **Address family** | IPv4 · IPv6 · dual-stack · neither | yes | `cidr:` and `address:` take both; `nat:` names which families are translated |
| **Interfaces per machine** | one · several | yes | `at:` takes a list |
| **Reachability policy, host** | symmetric · asymmetric | yes | `inbound:` — a routable machine that refuses everything |
| **Reachability policy, network** | open · segmented · asymmetric between segments | yes | `policy:` — inter-segment rules, as a segmented router enforces |
| **Segments per gateway** | one · several behind one router | yes | identical gateway declarations mean one gateway machine |
| **Gateway state** | permanent · expiring mappings | yes | `mapping_ttl:` |
| **Path MTU** | standard · reduced | yes | `segments[].mtu` |
| **Overlapping ranges** | distinct · two sites both on `192.168.1.0/24` | yes | segments may carry the same range |
| **Segment count** | one · many · isolated island | yes | `kind:` distinguishes an island from a public network |
| **Adjacency** | same segment · routed · unrelated public networks | yes | public segments are routed, never bridged — so no two are adjacent unless declared so |
| **Gateway depth** | direct · one gateway · nested | **partly** | `to:` chains, so nesting exists; `published:` names one gateway, so forwarding through two does not |
| **Address stability** | static · dynamic · changes mid-run | **partly** | a machine can be *moved*; an address changing under it in place cannot be stated |
| **Path quality** | latency · loss · bandwidth | **no**, deliberately | changes performance, not correctness — modelling it makes a network simulator, not a fixture |
### What closing the gaps changed
**Address family was not a field.** It changed the position model: a machine behind a household
gateway is typically *unforwardable on v4 and directly attached on v6, simultaneously*. So
reachability is a property of *(machine, family)*, and *"can these two nodes reach each other"*
is no longer a yes/no question — it is asked once per family, and the asymmetric answers are the
interesting ones. That distinction did not exist in the model an hour ago and would have been
discovered by a mesh failing over one family while reporting the other.
**`inbound:` became necessary because of v6.** With NAT, unreachability was implied by the
topology. With a globally routable v6 address, a machine is reachable unless something refuses —
so refusing has to be sayable, or v6 addressing would silently imply reachability.
**`nat:` became a list rather than a boolean** for the same reason: a real gateway translates v4
and routes v6, and a boolean cannot say that.
### What remains open, and whether it matters
Two partial axes, both extensible when something needs them, neither blocking: **nested
forwarding** and **an address changing in place**. A machine can already be moved, which covers
the roaming case; what is missing is a lease expiring underneath a machine that stays put.
One deliberate exclusion: **path quality**. Latency and loss change how fast the mesh is, not
whether it is correct. If a timeout turns out to be load-bearing that judgement should be
revisited — and it would be revisited by a real failure, which is the right trigger.
### The honest summary
The model now covers **where a machine sits**, **what the path between machines is like**, and
**what is permitted between them** — across both address families. Together those are what the
mesh's reachability logic turns on.
It contains almost no equipment, by the test above: a device that does not change what an IP
packet can do has no fault for a scenario to catch. What it does contain is every device that
does — which turns out to be the router, and a modem only when the modem is also a router.
What it still does not model is *change over time* beyond moving a machine, *degradation* short
of failure, and *publishing through two gateways at once*. The first two are chosen; the third
is the one real absence, and it is exactly the double-NAT case.
## Open
- **`user` and `edge` profiles have no scenario.** A lab machine is always privileged, so the
two profiles that exist for unprivileged and phone-like participation cannot be exercised.
Either the lab grows a way to run the host unprivileged, or those profiles are developed
against something that is not a virtual machine. This is the largest gap.
- **Where `place:` gets its artifacts from.** Before the mesh is self-hosting these come from
outside; afterwards from the mesh itself. The declaration should not have to care, which
suggests a named source rather than a path.
- **Nested forwarding** — `published:` names one gateway, so a machine behind two cannot be
published through both.
- **An address changing in place**, as a DHCP lease expiring under a machine that has not moved.