The full topology now raises: four machines, three routers, a transit router, six segments, in 35 seconds. Everything the declaration model can express except `place`, which is refused because the node host it would place does not exist yet. Transit was a real gap, not a bug. The design says public networks are unrelated and routed to each other, never bridged — and I built the segments and never built the thing that routes between them, so three public networks were islands and nothing crossed. A transit router now holds an interface on every public segment, forwarding and no translation: the closest thing the lab has to the internet, deliberately dumb. Proven rather than asserted, by ping TTL across the raised topology: within one segment ttl=64 no hops across two unrelated public networks ttl=62 gateway + transit multicast between public networks 0 replies A flat internet would have shown ttl=64 and answered multicast — which would let a node discover a peer it could never reach in production, and report success. That is the fault the as-is layer records the mesh already hitting with multicast name resolution. inbound: deny is implemented as a host firewall on the machine, read back after applying. A declared refusal that silently did not load leaves the machine wide open, which looks exactly like a machine that is working. Established and related traffic is accepted, so a defended machine can still dial out rather than being a disconnected one. Verified by running, all of it: home -> devices (policy allow) reachable devices -> home (policy deny) blocked behind unforwardable NAT -> out reachable in -> behind unforwardable NAT unreachable inbound: deny, dialling out reachable reaching a machine that denies inbound refused The two routers differ exactly as declared: the forwardable one carries the policy rule and no inbound drop, the unforwardable one carries `ct state new drop` and no DNAT.
8.3 KiB
mesh-lab
The lab: a disposable Novox Mesh on one machine.
It ships to nobody. It runs on a workstation, raises virtual machines, puts things inside them, and throws them away.
Why it exists first
The node host takes over a machine's packages, services and network. It cannot be developed against a machine anyone needs — so the place to develop it has to exist before it does.
That makes this repository phase 0 of the migration, ahead of every tier it will later test.
Two classes of scenario
| Bootstrap | Full | |
|---|---|---|
| Contains | virtual machines, the node host, a pinned substrate bundle | a complete mesh: forge, control plane, delivery, modules |
| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify |
| Exercises | tiers 0 and 1 | tier 2 and above, and modules |
| Exists to | develop the mesh | test what runs on it |
The bootstrap scenario is a strict subset — same virtualisation, same networking, same lifecycle, stopping before a control plane exists. The full scenario is reached by putting more inside the machines, not by building a second thing.
Using it
mesh-lab check can this machine run scenarios at all
mesh-lab validate scenarios/x.yml parse and check, raising nothing
mesh-lab raise scenarios/x.yml materialise it, wait until the machines are USABLE
mesh-lab list instances currently standing
mesh-lab exec <instance> <machine> -- <cmd...>
mesh-lab snapshot <instance> <label>
mesh-lab restore <instance> <label>
mesh-lab destroy <instance>
check refuses rather than warns. A machine without copy-on-write storage runs scenarios
correctly and snapshots roughly 76× slower — which does not make the lab slow, it makes it
unused, and a warning about that is read once and ignored forever.
If the incus socket is not reachable as your user — the group was granted to a session that
already existed — set MESH_LAB_INCUS="sudo -n incus".
What a scenario declares
The underlay: what a hosting provider and a home router would provide, and nothing the mesh is responsible for.
segments:
hosting: # one public network
kind: public
cidr: [192.0.2.0/24, "2001:db8:a::/48"]
isp-home: # another, unrelated — routed to it, never bridged
kind: public
cidr: [198.51.100.0/24, "2001:db8:b::/48"]
home:
kind: private
cidr: [192.168.1.0/24, "2001:db8:b:1::/64"]
mtu: 1492
gateway:
to: isp-home
address: [198.51.100.7] # what the world sees this network as
nat: [v4] # v4 translated, v6 routed
forwardable: true
mapping_ttl: 120s
machines:
home-server:
at: { segment: home, address: [192.168.1.135, "2001:db8:b:1::135"] }
published: [{ port: 443, on: home }]
inbound: allow
It declares nothing about overlay addresses, hubs, peering, names or certificates. Those are what the mesh does, and a scenario that supplied them would be certifying its own work.
Public segments must use documentation ranges (RFC 5737, RFC 3849) and the validator refuses anything else before raising. That is not pedantry: the mesh decides public-versus-private by matching the address, so a private range on a segment meant to be routable makes the mesh silently never form — no error, nothing to notice.
Reaching in
Everything goes through incus, never over IP. A scenario is a closed address space, so two instances raised from one declaration hold the same addresses and never meet — and the workstation has no route into either.
So a reachability question is asked from inside: can this machine reach that one is
exec on the first, testing the second. The workstation's opinion would be a different
question with a misleadingly similar answer.
What is implemented, and what is not
The declaration model is complete — it is the design's shape, and validating against it is useful before any of it can be raised. The runtime is not, and the gap is refused rather than ignored:
| segments as isolated links | works |
| machines, multi-homed or detached | works |
| declared addresses, both families | works |
| segment MTU | works |
| raise · exec · snapshot · restore · destroy · list | works |
| gateways, NAT, masquerade | works |
published: ports (DNAT through the gateway's address) |
works |
mapping_ttl: (conntrack timeout) |
works, and verified after setting — a declared expiry that silently did not apply would be the fault this catches |
forwardable: false |
works — outbound only, no DNAT, unsolicited inbound dropped |
policy: between segments |
works, asymmetric |
inbound: deny |
works — host firewall, read back after applying |
| several public networks, routed not bridged | works — a transit router, never a shared bridge |
place: |
refused at raise — the node host it would place does not exist yet |
raise refuses a scenario declaring anything in the lower half, naming every gap. It does not
raise a mesh that silently lacks what it declared — that is the fault this lab exists to catch
(novox/hq 04-ISSUES/003: a firewall key declared in five manifests and read by no code, so a
manifest appears to restrict a port and restricts nothing).
the-ordinary-shape.yml therefore validates and does not raise. That is the intended state:
it is the topology being built toward, and the tool says exactly what is missing.
Measured on a workstation
| one machine | two machines | two machines + a router | |
|---|---|---|---|
| raise, to usable | 12.5 s | 14.6 s | 32 s |
| snapshot | 0.14 s | 0.28 s | — |
| restore, to usable again | 10.5 s | 11.6 s | — |
A router adds seconds, not a boot: it is a container, because it is scenery rather than
something under test (novox/hq ADR 0033).
Verified by running, not asserted — a machine at 192.168.1.135 behind a household
gateway, reached from a machine on a routable address:
home-server -> anchor 0% loss, through masquerade
anchor -> 192.168.1.135 (private, direct) unreachable ✓
anchor -> 192.0.2.50:8080 (the GATEWAY) HTTP 200
home -> devices (policy allow) reachable ✓
devices -> home (policy deny) blocked ✓
roamer behind unforwardable NAT -> anchor reachable ✓ (outbound only)
anchor -> roamer unreachable ✓
workstation with inbound: deny, dialling out reachable ✓ (defended, not disconnected)
home-server -> workstation refused ✓
The third line is the case research 004 says only exists in production.
Routed, never bridged, proven rather than asserted — ping TTL across the full topology:
within one segment ttl=64 no hops
across two unrelated public networks ttl=62 gateway + transit
multicast between public networks 0 replies
A flat "internet" would have shown ttl=64 and answered multicast, which would have let a node discover a peer it could never reach in production — and report success.
Machines boot concurrently, so a second machine costs seconds rather than doubling the wait. Nearly all of the remaining time is boot, which cannot be avoided.
These numbers depend entirely on a copy-on-write pool. On dir the same snapshot takes 9.9 s
and a full copy of the disk, and a second one did not finish in two minutes — which is why
check refuses rather than warns.
Where the reasoning lives
Design and decisions are in novox/hq, not here:
03-DESIGN/01-to-be/02-scenario-declaration.md— what a scenario declares03-DESIGN/01-to-be/03-scenario-lifecycle.md— what happens to one02-DECISIONS/0031-the-lab-provides-the-underlay.md02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md
This repository carries implementation. It does not carry decisions.
Development
No build step — Node strips the types.
npm test the declaration layer, offline
npm run typecheck
The lifecycle is not unit-tested. It talks to a hypervisor, and a fake one would assert that the fake behaves as expected — which is the shape of test this project exists to stop shipping. It is exercised by raising real scenarios.