Commit Graph
4 Commits
Author SHA1 Message Date
jschoubben a6b7d67e19 Gateways sharing an address are one gateway
Found by asking what gw-devices and gw-home actually were, in a picture that
finally made them easy to see side by side.

planRouters grouped on the exact address list, so `home` declaring a v4 and a v6
address and `devices` declaring only the v4 became two router containers — both
holding 198.51.100.7 on the same segment. The lab raised it without complaint.

Not theoretical. On the raised instance the transit router resolved that one
address to two different MACs across a cache flush:

    198.51.100.7 -> 02:c9:16:70:23:29   (gw0, which HAS the :443 dnat)
    198.51.100.7 -> 02:bd:75:0b:b0:75   (gw1, which has none)

So home-server's published port worked or did not depending on which container
answered ARP last — intermittent, and it would have presented as a flaky test
rather than as a broken scenario.

One public address is one box. Checked against the thing this models rather than
argued from the model: a bridged modem, a single gateway holding the public
address, one network behind it, and every port forward landing on one host at
that address. Two routers on one address is not a topology, it is a collision.

Gateways to the same segment sharing any address are now one router and their
address lists union, so a v6 address declared on only one of the segments it
serves is still carried. Where such declarations disagree on nat, forwardable or
mapping_ttl, validate refuses — one box cannot behave two ways.

the-ordinary-shape now raises 7 machines instead of 8, and gw0 holds the public
address on eth0 while serving home on eth1 and devices on eth2.
2026-08-24 23:43:23 +02:00
jschoubben 2243618f01 Draw a scenario, from the declaration and from the hypervisor
`mesh-lab diagram` renders a scenario as draw.io, from either source, through
one layout — so a difference between what was asked for and what exists is a
difference you can see.

The shape says what a resource is and is fixed per kind. The badges say what is
true about that particular one and come entirely from metadata: translation,
forwardability, mapping expiry, refuses-inbound, container-or-VM, running. The
interesting properties of a network are exactly the ones with no visual
consequence — a translated address looks identical to an untranslated one.

For the live picture to be a record rather than a restatement, raise now writes
down what it applied: a segment's kind, ranges and MTU on the link; a gateway's
translation, forwardability and expiry on the gateway; inbound: deny on the
machine. Every behavioural tag is written AFTER the thing works, never at
creation — a failed raise leaves wreckage standing on purpose, and a picture of
that wreckage must not badge translation the router never got.

The pairing earned itself immediately: drawn side by side, every virtual machine
held no addresses. A container's interface carries the device's name and a VM
names its own, so joining them by name silently dropped one whole class of
machine. Fixed by joining on MAC.

Also brings tests under the typecheck gate, which caught integration timeouts
being passed as a 4th argument and therefore ignored entirely.
2026-08-24 22:53:00 +02:00
jschoubben 5d01006eab Transit, host firewalls, and the whole topology raising
The full topology now raises: four machines, three routers, a transit
router, six segments, in 35 seconds. Everything the declaration model can
express except `place`, which is refused because the node host it would
place does not exist yet.

Transit was a real gap, not a bug. The design says public networks are
unrelated and routed to each other, never bridged — and I built the
segments and never built the thing that routes between them, so three
public networks were islands and nothing crossed. A transit router now
holds an interface on every public segment, forwarding and no translation:
the closest thing the lab has to the internet, deliberately dumb.

Proven rather than asserted, by ping TTL across the raised topology:

  within one segment                     ttl=64   no hops
  across two unrelated public networks   ttl=62   gateway + transit
  multicast between public networks      0 replies

A flat internet would have shown ttl=64 and answered multicast — which
would let a node discover a peer it could never reach in production, and
report success. That is the fault the as-is layer records the mesh already
hitting with multicast name resolution.

inbound: deny is implemented as a host firewall on the machine, read back
after applying. A declared refusal that silently did not load leaves the
machine wide open, which looks exactly like a machine that is working.
Established and related traffic is accepted, so a defended machine can
still dial out rather than being a disconnected one.

Verified by running, all of it:

  home -> devices (policy allow)               reachable
  devices -> home (policy deny)                blocked
  behind unforwardable NAT -> out              reachable
  in -> behind unforwardable NAT               unreachable
  inbound: deny, dialling out                  reachable
  reaching a machine that denies inbound       refused

The two routers differ exactly as declared: the forwardable one carries the
policy rule and no inbound drop, the unforwardable one carries `ct state
new drop` and no DNAT.
2026-08-24 01:49:30 +02:00
jschoubben a270cd5b02 Routers: NAT, port forwarding, policy and mapping expiry
A gateway is the one implicit machine in a declaration — a scenario says a
segment sits behind one and never names the thing that serves it. This
materialises it.

A router is a container, not a virtual machine, because it is scenery
rather than something under test (hq ADR 0033). Verified before building
that a plain unprivileged container can do all of it: ip_forward and ipv6
forwarding settable, nftables masquerade accepted, and the conntrack
timeouts mapping_ttl depends on both writable. No privileged mode.

Verified by running, on a machine behind a household gateway reached from
one on a routable address:

  home-server -> anchor                      0% loss, through masquerade
  anchor -> 192.168.1.135 (private, direct)  unreachable
  anchor -> 192.0.2.50:8080 (the GATEWAY)    HTTP 200

The last line is the published-but-behind-NAT case research 004 says only
exists in production. It is now a 32-second scenario on a workstation.

Segments sharing a gateway declaration share ONE router — that is what a
VLAN-capable router is, and two routers sharing an external address would
not work anyway.

mapping_ttl is read back after setting rather than assumed. Those sysctls
are not on every kernel, and a scenario that declared an expiring mapping
and silently got a permanent one would be exactly the fault being built
against.

Four bugs found by running it, three of them the same fault — a failure
made invisible.

The router had no route to a package repository, by design, so installing
nftables at raise time could not work. The image is now built once with
temporary connectivity and cached; every scenario after that needs no
network. That failure was hidden behind `|| true`, which is why it took a
raise to find.

The builder then failed on DNS: exec works before a container has an
address, and I had treated usable as ready. It now waits for the thing
actually needed.

The stock Alpine image ships `auto eth0 / inet dhcp` and its boot-time
networking service flushed the static address the scenario set — on eth0
only, so the outside interface came up bare while inside ones were fine.
The image build now neutralises it: a router reconfiguring itself from an
image default is the lab overriding the declaration. `ip addr add … || true`
had hidden this too, and is now `ip addr replace` with no swallow.

And routers were orphaned by destroy, holding their networks open so
destroy reported removing zero segments. They now carry the same machine
tag as everything else, so one query finds an instance's resources.
2026-08-24 01:37:19 +02:00