a270cd5b02340ed58528e046c026e49b6ab5f2dc
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a270cd5b02 |
Routers: NAT, port forwarding, policy and mapping expiry
A gateway is the one implicit machine in a declaration — a scenario says a segment sits behind one and never names the thing that serves it. This materialises it. A router is a container, not a virtual machine, because it is scenery rather than something under test (hq ADR 0033). Verified before building that a plain unprivileged container can do all of it: ip_forward and ipv6 forwarding settable, nftables masquerade accepted, and the conntrack timeouts mapping_ttl depends on both writable. No privileged mode. Verified by running, on a machine behind a household gateway reached from one on a routable address: home-server -> anchor 0% loss, through masquerade anchor -> 192.168.1.135 (private, direct) unreachable anchor -> 192.0.2.50:8080 (the GATEWAY) HTTP 200 The last line is the published-but-behind-NAT case research 004 says only exists in production. It is now a 32-second scenario on a workstation. Segments sharing a gateway declaration share ONE router — that is what a VLAN-capable router is, and two routers sharing an external address would not work anyway. mapping_ttl is read back after setting rather than assumed. Those sysctls are not on every kernel, and a scenario that declared an expiring mapping and silently got a permanent one would be exactly the fault being built against. Four bugs found by running it, three of them the same fault — a failure made invisible. The router had no route to a package repository, by design, so installing nftables at raise time could not work. The image is now built once with temporary connectivity and cached; every scenario after that needs no network. That failure was hidden behind `|| true`, which is why it took a raise to find. The builder then failed on DNS: exec works before a container has an address, and I had treated usable as ready. It now waits for the thing actually needed. The stock Alpine image ships `auto eth0 / inet dhcp` and its boot-time networking service flushed the static address the scenario set — on eth0 only, so the outside interface came up bare while inside ones were fine. The image build now neutralises it: a router reconfiguring itself from an image default is the lab overriding the declaration. `ip addr add … || true` had hidden this too, and is now `ip addr replace` with no swallow. And routers were orphaned by destroy, holding their networks open so destroy reported removing zero segments. They now carry the same machine tag as everything else, so one query finds an instance's resources. |
||
|
|
a27d861d3b |
Scenario lifecycle: raise, exec, snapshot, restore, destroy
A declaration goes in and a disposable mesh comes out. Verified on a workstation, not asserted: two machines raised and addressed in 14.6s, snapshot 0.28s, restore-to-usable 11.6s, both families pinging with no loss, and the workstation with no route into any of it. The declaration layer implements the model in full — three positions a machine can be in, keyed on forwardability; gateways carrying the address the world sees them as; both address families; multi-homing; MTU; inter-segment policy. It is validated hard because the failures it prevents are silent: a private range on a public segment produces no error, the mesh simply never forms. Public segments are refused unless they use RFC 5737 or RFC 3849 space, and a range wider than the reserved block is refused too. 33 tests, all offline. The runtime implements less than the model, and refuses the difference. A scenario declaring gateways, published ports, policy, inbound deny or place is rejected at raise with every gap named. Raising it would produce a mesh that silently lacks what it declared, which is the fault this lab exists to catch — 04-ISSUES/003, where a firewall key is declared in five manifests and read by no code. Three bugs found by review and by running it, all of one family: The readiness check truthiness-tested incusOk's return. `exec … true` succeeds with EMPTY output, so every machine reported unreachable while incus exec on it worked perfectly. succeeds() now exists so the mistake is not available, and network delete had the same bug — it counted zero segments removed while removing them. list() split instance from machine on the last dash, so a machine called home-server absorbed half the instance id and destroy found nothing. Resources are now found by the metadata they carry, never by name. restore reported success in 0.79s while the machine's agent was still starting, so the next command failed. Both raise and restore now wait for usable and say how long that took — reporting the earlier number is transport reported as effect, which is the fault the lab is being built to find. Two incus behaviours worth recording. Its CLI reads a YAML definition from stdin when stdin is not a terminal, so a spawned command hangs until the timeout kills it and arrives with empty stderr — a failure with no explanation, on a command that works when typed. And it assigns a MAC at runtime without recording it in device config, so MACs are derived and set explicitly, which the guest needs anyway: it names interfaces by bus position, and matching by name configures the wrong one on a multi-homed machine. No build step; Node strips the types. The lifecycle has no unit tests because a fake hypervisor would assert that the fake behaves as expected, which is the shape of test this project exists to stop shipping. |