Reviewed and the criticism was right: 1,072 of 2,128 lines untested, all of it the half that touches the hypervisor, and no gate. The verification I had done was real — pings across NAT, TTL counts, ruleset comparisons — and none of it survived the terminal it ran in, which is 04-ISSUES/005 in miniature. Ten integration tests against a real hypervisor, each named for what it defends. ADR 0031: a raised machine carries no overlay, no wireguard, no mesh config — a scenario that pre-built peering would certify its own work. ADR 0032: exec is the only way in. ADR 0033: routers are containers while machines are virtual machines. And the design's claims: raise waits for usable, snapshots are whole-scenario, NAT hides a private address, published reaches the machine at the gateway's address. Mocking the hypervisor is forbidden, so they skip with a reason on a machine that cannot raise scenarios rather than passing green having checked nothing. The suite earned itself on its first run. It found that a snapshot of a running machine could miss a file written seconds earlier — not stale, absent — because the write was still in the guest's page cache. That is exactly the question the lifecycle design listed as open: does a scenario snapshot need the machines stopped? It does not, but it does need them flushed. snapshot now syncs every machine before capturing, and the design records the answer. The fix buys write-durability, not application-consistency: anything mid-transaction is still captured mid-transaction, and that is now stated rather than assumed. npm run check is the gate — typecheck, 40 unit tests, 10 integration tests.
221 lines
9.6 KiB
Markdown
221 lines
9.6 KiB
Markdown
# mesh-lab
|
||
|
||
The lab: a disposable Novox Mesh on one machine.
|
||
|
||
It ships to nobody. It runs on a workstation, raises virtual machines, puts things inside
|
||
them, and throws them away.
|
||
|
||
## Why it exists first
|
||
|
||
The node host takes over a machine's packages, services and network. It cannot be developed
|
||
against a machine anyone needs — so the place to develop it has to exist before it does.
|
||
|
||
That makes this repository **phase 0** of the migration, ahead of every tier it will later
|
||
test.
|
||
|
||
## Two classes of scenario
|
||
|
||
| | **Bootstrap** | **Full** |
|
||
|---|---|---|
|
||
| Contains | virtual machines, the node host, a pinned substrate bundle | a complete mesh: forge, control plane, delivery, modules |
|
||
| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify |
|
||
| Exercises | tiers 0 and 1 | tier 2 and above, and modules |
|
||
| Exists to | **develop the mesh** | **test what runs on it** |
|
||
|
||
The bootstrap scenario is a **strict subset** — same virtualisation, same networking, same
|
||
lifecycle, stopping before a control plane exists. The full scenario is reached by putting more
|
||
inside the machines, not by building a second thing.
|
||
|
||
## Using it
|
||
|
||
```
|
||
mesh-lab check can this machine run scenarios at all
|
||
mesh-lab validate scenarios/x.yml parse and check, raising nothing
|
||
mesh-lab raise scenarios/x.yml materialise it, wait until the machines are USABLE
|
||
mesh-lab list instances currently standing
|
||
mesh-lab exec <instance> <machine> -- <cmd...>
|
||
mesh-lab snapshot <instance> <label>
|
||
mesh-lab restore <instance> <label>
|
||
mesh-lab destroy <instance>
|
||
```
|
||
|
||
`check` refuses rather than warns. A machine without copy-on-write storage runs scenarios
|
||
correctly and snapshots roughly 76× slower — which does not make the lab slow, it makes it
|
||
unused, and a warning about that is read once and ignored forever.
|
||
|
||
If the incus socket is not reachable as your user — the group was granted to a session that
|
||
already existed — set `MESH_LAB_INCUS="sudo -n incus"`.
|
||
|
||
## What a scenario declares
|
||
|
||
The **underlay**: what a hosting provider and a home router would provide, and nothing the
|
||
mesh is responsible for.
|
||
|
||
```yaml
|
||
segments:
|
||
hosting: # one public network
|
||
kind: public
|
||
cidr: [192.0.2.0/24, "2001:db8:a::/48"]
|
||
isp-home: # another, unrelated — routed to it, never bridged
|
||
kind: public
|
||
cidr: [198.51.100.0/24, "2001:db8:b::/48"]
|
||
home:
|
||
kind: private
|
||
cidr: [192.168.1.0/24, "2001:db8:b:1::/64"]
|
||
mtu: 1492
|
||
gateway:
|
||
to: isp-home
|
||
address: [198.51.100.7] # what the world sees this network as
|
||
nat: [v4] # v4 translated, v6 routed
|
||
forwardable: true
|
||
mapping_ttl: 120s
|
||
machines:
|
||
home-server:
|
||
at: { segment: home, address: [192.168.1.135, "2001:db8:b:1::135"] }
|
||
published: [{ port: 443, on: home }]
|
||
inbound: allow
|
||
```
|
||
|
||
It declares **nothing** about overlay addresses, hubs, peering, names or certificates. Those
|
||
are what the mesh does, and a scenario that supplied them would be certifying its own work.
|
||
|
||
Public segments must use documentation ranges (RFC 5737, RFC 3849) and the validator refuses
|
||
anything else **before raising**. That is not pedantry: the mesh decides public-versus-private
|
||
by matching the address, so a private range on a segment meant to be routable makes the mesh
|
||
silently never form — no error, nothing to notice.
|
||
|
||
## Reaching in
|
||
|
||
Everything goes through incus, never over IP. A scenario is a closed address space, so two
|
||
instances raised from one declaration hold the same addresses and never meet — and the
|
||
workstation has no route into either.
|
||
|
||
So a reachability question is asked **from inside**: *can this machine reach that one* is
|
||
`exec` on the first, testing the second. The workstation's opinion would be a different
|
||
question with a misleadingly similar answer.
|
||
|
||
## What is implemented, and what is not
|
||
|
||
The declaration model is complete — it is the design's shape, and validating against it is
|
||
useful before any of it can be raised. **The runtime is not**, and the gap is refused rather
|
||
than ignored:
|
||
|
||
| | |
|
||
|---|---|
|
||
| segments as isolated links | **works** |
|
||
| machines, multi-homed or detached | **works** |
|
||
| declared addresses, both families | **works** |
|
||
| segment MTU | **works** |
|
||
| raise · exec · snapshot · restore · destroy · list | **works** |
|
||
| gateways, NAT, masquerade | **works** |
|
||
| `published:` ports (DNAT through the gateway's address) | **works** |
|
||
| `mapping_ttl:` (conntrack timeout) | **works**, and verified after setting — a declared expiry that silently did not apply would be the fault this catches |
|
||
| `forwardable: false` | **works** — outbound only, no DNAT, unsolicited inbound dropped |
|
||
| `policy:` between segments | **works**, asymmetric |
|
||
| `inbound: deny` | **works** — host firewall, read back after applying |
|
||
| several public networks, routed not bridged | **works** — a transit router, never a shared bridge |
|
||
| `place:` | **refused at raise** — the node host it would place does not exist yet |
|
||
|
||
`raise` refuses a scenario declaring anything in the lower half, naming every gap. It does not
|
||
raise a mesh that silently lacks what it declared — that is the fault this lab exists to catch
|
||
(`novox/hq` 04-ISSUES/003: a firewall key declared in five manifests and read by no code, so a
|
||
manifest appears to restrict a port and restricts nothing).
|
||
|
||
`the-ordinary-shape.yml` therefore validates and does not raise. That is the intended state:
|
||
it is the topology being built toward, and the tool says exactly what is missing.
|
||
|
||
## Measured on a workstation
|
||
|
||
| | one machine | two machines | two machines + a router |
|
||
|---|---|---|---|
|
||
| raise, to usable | 12.5 s | 14.6 s | 32 s |
|
||
| snapshot | 0.14 s | 0.28 s | — |
|
||
| restore, to usable again | 10.5 s | 11.6 s | — |
|
||
|
||
A router adds seconds, not a boot: it is a container, because it is scenery rather than
|
||
something under test (`novox/hq` ADR 0033).
|
||
|
||
**Verified by running**, not asserted — a machine at `192.168.1.135` behind a household
|
||
gateway, reached from a machine on a routable address:
|
||
|
||
```
|
||
home-server -> anchor 0% loss, through masquerade
|
||
anchor -> 192.168.1.135 (private, direct) unreachable ✓
|
||
anchor -> 192.0.2.50:8080 (the GATEWAY) HTTP 200
|
||
home -> devices (policy allow) reachable ✓
|
||
devices -> home (policy deny) blocked ✓
|
||
roamer behind unforwardable NAT -> anchor reachable ✓ (outbound only)
|
||
anchor -> roamer unreachable ✓
|
||
workstation with inbound: deny, dialling out reachable ✓ (defended, not disconnected)
|
||
home-server -> workstation refused ✓
|
||
```
|
||
|
||
The third line is the case research 004 says only exists in production.
|
||
|
||
**Routed, never bridged**, proven rather than asserted — ping TTL across the full topology:
|
||
|
||
```
|
||
within one segment ttl=64 no hops
|
||
across two unrelated public networks ttl=62 gateway + transit
|
||
multicast between public networks 0 replies
|
||
```
|
||
|
||
A flat "internet" would have shown ttl=64 and answered multicast, which would have let a node
|
||
discover a peer it could never reach in production — and report success.
|
||
|
||
Machines boot concurrently, so a second machine costs seconds rather than doubling the wait.
|
||
Nearly all of the remaining time is boot, which cannot be avoided.
|
||
|
||
These numbers depend entirely on a copy-on-write pool. On `dir` the same snapshot takes 9.9 s
|
||
and a full copy of the disk, and a second one did not finish in two minutes — which is why
|
||
`check` refuses rather than warns.
|
||
|
||
## Where the reasoning lives
|
||
|
||
Design and decisions are in [`novox/hq`](https://git.novox.be/novox/hq), not here:
|
||
|
||
- `03-DESIGN/01-to-be/02-scenario-declaration.md` — what a scenario declares
|
||
- `03-DESIGN/01-to-be/03-scenario-lifecycle.md` — what happens to one
|
||
- `02-DECISIONS/0031-the-lab-provides-the-underlay.md`
|
||
- `02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md`
|
||
|
||
This repository carries implementation. It does not carry decisions.
|
||
|
||
## Development
|
||
|
||
No build step — Node strips the types.
|
||
|
||
```
|
||
npm test the declaration layer, offline
|
||
npm run typecheck
|
||
```
|
||
|
||
```
|
||
npm test the declaration layer, offline, 40 tests
|
||
npm run test:integration real scenarios against a real hypervisor, 10 tests
|
||
npm run check typecheck + both — this is the gate
|
||
```
|
||
|
||
**A test names the decision it defends** (`novox/hq` ADR 0034). A decision with no test is one
|
||
that will quietly stop being true, and nobody learns that from a document:
|
||
|
||
| Test | Defends |
|
||
|---|---|
|
||
| the lab provides the underlay and nothing of the overlay | ADR 0031 |
|
||
| the workstation has no route into the scenario | ADR 0032 |
|
||
| a router is a container while machines are virtual machines | ADR 0033 |
|
||
| raise waits for *usable*, not for the call to return | the lifecycle design |
|
||
| snapshots are whole-scenario | the lifecycle design |
|
||
| a public range that is not documentation space is refused | the declaration design |
|
||
| a scenario declaring what cannot be materialised is refused | the declaration design |
|
||
|
||
**Mocking the hypervisor is forbidden.** A fake would assert that the fake behaves as expected,
|
||
which is the shape of test this project exists to stop shipping. Integration tests skip with a
|
||
reason on a machine that cannot raise scenarios, rather than passing green having checked
|
||
nothing.
|
||
|
||
That suite earned itself on its first run: it found that a snapshot of a running machine could
|
||
miss a file written seconds earlier — not stale, **absent** — because the write was still in
|
||
the guest's page cache. The design had listed that as an open question. The test answered it,
|
||
and `snapshot` now flushes first.
|