d56a0c8fef01327453ef68aa83573f3e7b76f800
5
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4097ff92c1 |
Repoint ADR references after HQ consolidated 65 records to 23
Comments naming records that no longer exist now point at the consolidated record holding their reasoning -- the four lab records are 0016, a test defends a decision is 0017. |
||
|
|
16c13807a9 |
place: the host — the lab acquires a consumer
The lab raised an underlay and put nothing on it: correct, and useless, because the thing it exists to test did not exist. Tier 0 now does, so `place: [host]` works and a raised scenario finally contains something. The refusal narrows rather than disappearing. A scenario placing a host and a substrate is told which half is missing, by name — not that `place:` is unsupported when half of it now works. Placement reads back rather than assuming. A file arriving is not a host working, so the binary is run before it is trusted to answer questions, and what it reports is read from the machine (ADR 0035). The binary comes from an explicit path, because the declaration design leaves where artifacts come from open and a search would harden into the answer by accident. The integration test that matters is the one asserting the host reports the MACHINE and not the workstation that placed it. A raised VM and this workstation differ in every capability — root versus uid 1000, a clean init versus a degraded one, no docker versus docker, no wireguard versus wg0 — so a host reporting the wrong machine is obvious here and invisible anywhere else. And the placed host independently confirms ADR 0031: overlay absent on a freshly raised machine. The underlay suite already asserted that by looking for wireguard interfaces; this is a second witness rather than the same check twice. Two tests failed the moment placement worked, which is what they were for. They defended "there is nothing to place yet" while that was true; the decision changed, so they change with it rather than being deleted. Gate: 75 unit, 20 integration. |
||
|
|
5d01006eab |
Transit, host firewalls, and the whole topology raising
The full topology now raises: four machines, three routers, a transit router, six segments, in 35 seconds. Everything the declaration model can express except `place`, which is refused because the node host it would place does not exist yet. Transit was a real gap, not a bug. The design says public networks are unrelated and routed to each other, never bridged — and I built the segments and never built the thing that routes between them, so three public networks were islands and nothing crossed. A transit router now holds an interface on every public segment, forwarding and no translation: the closest thing the lab has to the internet, deliberately dumb. Proven rather than asserted, by ping TTL across the raised topology: within one segment ttl=64 no hops across two unrelated public networks ttl=62 gateway + transit multicast between public networks 0 replies A flat internet would have shown ttl=64 and answered multicast — which would let a node discover a peer it could never reach in production, and report success. That is the fault the as-is layer records the mesh already hitting with multicast name resolution. inbound: deny is implemented as a host firewall on the machine, read back after applying. A declared refusal that silently did not load leaves the machine wide open, which looks exactly like a machine that is working. Established and related traffic is accepted, so a defended machine can still dial out rather than being a disconnected one. Verified by running, all of it: home -> devices (policy allow) reachable devices -> home (policy deny) blocked behind unforwardable NAT -> out reachable in -> behind unforwardable NAT unreachable inbound: deny, dialling out reachable reaching a machine that denies inbound refused The two routers differ exactly as declared: the forwardable one carries the policy rule and no inbound drop, the unforwardable one carries `ct state new drop` and no DNAT. |
||
|
|
a270cd5b02 |
Routers: NAT, port forwarding, policy and mapping expiry
A gateway is the one implicit machine in a declaration — a scenario says a segment sits behind one and never names the thing that serves it. This materialises it. A router is a container, not a virtual machine, because it is scenery rather than something under test (hq ADR 0033). Verified before building that a plain unprivileged container can do all of it: ip_forward and ipv6 forwarding settable, nftables masquerade accepted, and the conntrack timeouts mapping_ttl depends on both writable. No privileged mode. Verified by running, on a machine behind a household gateway reached from one on a routable address: home-server -> anchor 0% loss, through masquerade anchor -> 192.168.1.135 (private, direct) unreachable anchor -> 192.0.2.50:8080 (the GATEWAY) HTTP 200 The last line is the published-but-behind-NAT case research 004 says only exists in production. It is now a 32-second scenario on a workstation. Segments sharing a gateway declaration share ONE router — that is what a VLAN-capable router is, and two routers sharing an external address would not work anyway. mapping_ttl is read back after setting rather than assumed. Those sysctls are not on every kernel, and a scenario that declared an expiring mapping and silently got a permanent one would be exactly the fault being built against. Four bugs found by running it, three of them the same fault — a failure made invisible. The router had no route to a package repository, by design, so installing nftables at raise time could not work. The image is now built once with temporary connectivity and cached; every scenario after that needs no network. That failure was hidden behind `|| true`, which is why it took a raise to find. The builder then failed on DNS: exec works before a container has an address, and I had treated usable as ready. It now waits for the thing actually needed. The stock Alpine image ships `auto eth0 / inet dhcp` and its boot-time networking service flushed the static address the scenario set — on eth0 only, so the outside interface came up bare while inside ones were fine. The image build now neutralises it: a router reconfiguring itself from an image default is the lab overriding the declaration. `ip addr add … || true` had hidden this too, and is now `ip addr replace` with no swallow. And routers were orphaned by destroy, holding their networks open so destroy reported removing zero segments. They now carry the same machine tag as everything else, so one query finds an instance's resources. |
||
|
|
a27d861d3b |
Scenario lifecycle: raise, exec, snapshot, restore, destroy
A declaration goes in and a disposable mesh comes out. Verified on a workstation, not asserted: two machines raised and addressed in 14.6s, snapshot 0.28s, restore-to-usable 11.6s, both families pinging with no loss, and the workstation with no route into any of it. The declaration layer implements the model in full — three positions a machine can be in, keyed on forwardability; gateways carrying the address the world sees them as; both address families; multi-homing; MTU; inter-segment policy. It is validated hard because the failures it prevents are silent: a private range on a public segment produces no error, the mesh simply never forms. Public segments are refused unless they use RFC 5737 or RFC 3849 space, and a range wider than the reserved block is refused too. 33 tests, all offline. The runtime implements less than the model, and refuses the difference. A scenario declaring gateways, published ports, policy, inbound deny or place is rejected at raise with every gap named. Raising it would produce a mesh that silently lacks what it declared, which is the fault this lab exists to catch — 04-ISSUES/003, where a firewall key is declared in five manifests and read by no code. Three bugs found by review and by running it, all of one family: The readiness check truthiness-tested incusOk's return. `exec … true` succeeds with EMPTY output, so every machine reported unreachable while incus exec on it worked perfectly. succeeds() now exists so the mistake is not available, and network delete had the same bug — it counted zero segments removed while removing them. list() split instance from machine on the last dash, so a machine called home-server absorbed half the instance id and destroy found nothing. Resources are now found by the metadata they carry, never by name. restore reported success in 0.79s while the machine's agent was still starting, so the next command failed. Both raise and restore now wait for usable and say how long that took — reporting the earlier number is transport reported as effect, which is the fault the lab is being built to find. Two incus behaviours worth recording. Its CLI reads a YAML definition from stdin when stdin is not a terminal, so a spawned command hangs until the timeout kills it and arrives with empty stderr — a failure with no explanation, on a command that works when typed. And it assigns a MAC at runtime without recording it in device config, so MACs are derived and set explicitly, which the guest needs anyway: it names interfaces by bus position, and matching by name configures the wrong one on a multi-homed machine. No build step; Node strips the types. The lifecycle has no unit tests because a fake hypervisor would assert that the fake behaves as expected, which is the shape of test this project exists to stop shipping. |