whole-mesh-full: the real segmented topology, and the overlay proven across the access point

Rewrite the flat three-node whole-mesh-full (separate anchor, one public segment)
into production's real shape: two segments and one access point. novox sits on
the routable `hosting` segment and IS the anchor — it runs the substrate, its own
service set, the overlay hub and public ingress; there is no separate anchor node.
ace, shanks and g14 sit on the household `home` segment behind a NAT gateway,
reachable from outside only through what they dial out to.

The bed drives, and verifies, the thing the flat beds never could: the WireGuard
overlay forming ACROSS the access point — a home node dialling novox's public hub
endpoint out through the gateway's masquerade, the handshake completing through the
NAT, the keepalive holding the hole open. Phase A proves it (handshake state + a
ping over the overlay) before any heavy module lands; Phase B converges both server
sets. With MESH_LAB_KEEP the instance is raised under a fixed id and left standing.

Collapsing the substrate onto novox exposed real facts the separate-anchor beds
never hit, fixed here:
- the substrate bundle advertises the broker at 192.0.2.10 (the old anchor); a
  token carries that verbatim as the endpoint a node dials, so with the substrate
  on novox it must be novox's own public address. Rewritten at apply (the cert is
  fingerprint-pinned, not hostname-checked, so only the address needs correcting).
- the two provider host-port collisions with the co-located substrate: postgres
  5432 vs the store's 127.0.0.1:5432, lavinmq 5672 vs the broker's 127.0.0.1:5672.
  Both provider host publishes are remapped off the substrate's ports.

And a lab limitation this first large-union bed exposed: the image registry VM took
the profile's default `dir` pool and a ~10GiB root, which the ~28GiB union of both
server sets overflows ("no space left on device"). raiseRegistry now places the
registry on the scenario's copy-on-write pool with a sized (default 80GiB, thin)
root disk, MESH_LAB_REGISTRY_DISK overridable.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
This commit is contained in:
2026-09-09 11:17:00 +02:00
parent a5f53f524a
commit 80b0670ebe
4 changed files with 345 additions and 199 deletions
+93 -32
View File
@@ -1,54 +1,104 @@
# The FULL mesh: both server sets on ONE substrate, converging together — the final stage of the
# whole-mesh rehearsal (novox/hq). Combines scenarios/whole-mesh-novox.yml and whole-mesh-ace.yml.
# The FULL mesh in its REAL production shape: two segments, one access point, one overlay.
#
# anchor — substrate ONLY (store, broker, control).
# novox — the 17-module novox set (providers + web apps + route-proxy + mailu + firewall).
# ace — the 24-module ace set (media/home stack), its /services/media library pre-created.
# This is the first multi-segment whole-mesh bed. The earlier flat whole-mesh-full sat every node
# on one public segment with a SEPARATE `anchor` carrying the substrate. Production is not flat, and
# there is no separate anchor: `novox` IS the anchor. It sits on the routable `hosting` segment,
# runs the substrate (store, broker, control) AND its own service set AND is the overlay hub and the
# public ingress. `ace`, `shanks` and `g14` sit on the household `home` segment BEHIND a NAT gateway
# — the access point — reachable from the outside only through what they dial out to.
#
# An overlay is placed across all three so cross-node `at` resolves. Each service node is
# self-contained (its own postgres/redis), so nothing crosses a node boundary except enrolment and
# the shared broker/store on anchor — which is exactly what this stage proves converges for two
# independent node-plans at once on one substrate.
# hosting (public, routable) home (private, behind the access point)
# novox 192.0.2.20 ── anchor ace 192.168.1.10 home server, media/IoT set
# substrate + novox set shanks 192.168.1.20 workstation (light)
# overlay hub, ingress g14 192.168.1.30 workstation (light)
#
# The `home` gateway masquerades v4 outbound and forwards inbound (an ordinary household router).
# Home nodes reach novox's public 192.0.2.20 by dialling OUT through it: the substrate broker (5671),
# the registry, and — the thing this bed exists to prove — the WireGuard overlay hub (51820/udp).
# The hub keepalive holds the NAT hole open so the tunnel, once formed, stays up. novox cannot
# initiate to a home node at all; every home↔novox path is either the overlay or a forwarded port.
#
# THE UNPROVEN THING (what the flat beds never tested): does the overlay tunnel FORM across the
# access point — a home node dialling novox's public hub endpoint, the handshake completing through
# the gateway's masquerade? The driving test verifies the WireGuard handshake and cross-segment
# reachability over the overlay explicitly, and reports form-vs-break as its headline.
#
# Substrate-on-novox collides on two host ports the separate-anchor beds never hit: the substrate
# store binds 127.0.0.1:5432 and novox's postgres provider publishes 5432; the substrate broker binds
# 5671 + 127.0.0.1:5672 and novox's lavinmq provider publishes 5672. The driving test REMAPS those two
# provider host publishes off the substrate's ports (consumers reach the providers over the mesh
# network on the container port, so the host side is free to move). Reported as a topology finding.
#
# MESH_LAB_HOST_BINARY=.../mesh-host MESH_LAB_BUNDLE=.../examples/substrate-first-node.lock
# The images are the UNION of the two per-server scenarios; every one is already built/pulled by the
# per-server bed prerequisites (scripts/build-module-runtime.sh, build-route-proxy-image.sh, the
# mailu/keycloak and media :mesh digest pulls).
# The images are the UNION of the novox set (feat/novox-conversions @ 431310f: the slug + roundcube
# fixes, so only-office/de-spiegel/amqp-email-forwarder now resolve) and the ace media/home set. Every
# one is already built/pulled by the per-server bed prerequisites.
scenario: whole-mesh-full
segments:
# The routable segment. novox lives here; the lab raises the image registry here too (a public
# IPv4 segment is what serves the images), and the overlay hub endpoint is a public address here.
hosting:
kind: public
cidr: [192.0.2.0/24]
# The household segment behind the access point. Its gateway is an ordinary home router: it
# masquerades v4 outbound, forwards inbound, and expires idle mappings after two minutes — which
# is exactly the NAT hole a WireGuard keepalive has to hold open.
home:
kind: private
cidr: [192.168.1.0/24]
gateway:
to: hosting
address: [192.0.2.50] # what the world sees the household as
nat: [v4]
forwardable: true
mapping_ttl: 120s
machines:
anchor:
at: { segment: hosting, address: [192.0.2.10] }
inbound: allow
memory: 4GiB
cpus: 4
disk: 20GiB
# The anchor: substrate (store, broker, control) + the whole novox service set + overlay hub +
# public ingress. Bigger than the flat bed's novox, because it now carries the substrate too.
novox:
at: { segment: hosting, address: [192.0.2.20] }
inbound: allow
memory: 16GiB
cpus: 6
disk: 100GiB
memory: 24GiB
cpus: 8
disk: 130GiB
# The home server: the whole ace media/home set — 24 modules, ~50 containers, several heavy
# (Plex, Home Assistant, Letta, Baserow, the UniFi JVM, mssql). Behind the gateway.
ace:
at: { segment: hosting, address: [192.0.2.30] }
at: { segment: home, address: [192.168.1.10] }
inbound: allow
memory: 18GiB
cpus: 6
disk: 120GiB
# Two workstations on the same home LAN. Light on purpose: they enrol, join the overlay, and run
# one small module (portainer) so a real module converges on each without heavy load. Same-LAN
# nodes with no overlay endpoint of their own hairpin the hub rather than peering directly, which
# is the normal case and is fine.
shanks:
at: { segment: home, address: [192.168.1.20] }
inbound: allow
memory: 3GiB
cpus: 2
disk: 30GiB
g14:
at: { segment: home, address: [192.168.1.30] }
inbound: allow
memory: 3GiB
cpus: 2
disk: 30GiB
images:
# --- substrate + shared ---
# --- substrate + shared (novox & ace both run postgres/redis/mssql/portainer) ---
- postgres:17-alpine
- cloudamqp/lavinmq:latest
- mesh-control:development
- redis:7-alpine
- portainer/portainer-ce:latest
- mcr.microsoft.com/mssql/server:2022-latest
- portainer/portainer-ce:latest
# --- novox server images ---
- minio/minio:latest
- mongo:7
@@ -61,13 +111,24 @@ images:
- registry:2
- registry-api.novox.be/novox/invoicing-app:latest
- registry-api.novox.be/novox/invoicing-api:latest
- ghcr.io/mailu/unbound:mesh
- ghcr.io/mailu/admin:mesh
- ghcr.io/mailu/dovecot:mesh
- ghcr.io/mailu/postfix:mesh
- ghcr.io/mailu/rspamd:mesh
- ghcr.io/mailu/webmail:mesh
- ghcr.io/mailu/nginx:mesh
- onlyoffice/documentserver:mesh
- registry-api.novox.be/novox/de-spiegel:latest
- registry-api.novox.be/novox/www:latest
- registry-api.novox.be/novox/amqp-email-forwarder:latest
- registry-api.novox.be/novox/photos-server:latest
- registry-api.novox.be/novox/photos-admin-client:latest
- registry-api.novox.be/novox/photos-client:latest
# The full Mailu 1.9 stack (mailu-redis reuses redis:7-alpine above).
- ghcr.io/mailu/unbound:1.9
- ghcr.io/mailu/admin:1.9
- ghcr.io/mailu/dovecot:1.9
- ghcr.io/mailu/postfix:1.9
- ghcr.io/mailu/rspamd:1.9
- ghcr.io/mailu/clamav:1.9
- ghcr.io/mailu/roundcube:1.9
- ghcr.io/mailu/radicale:1.9
- ghcr.io/mailu/fetchmail:1.9
- ghcr.io/mailu/nginx:1.9
# --- ace server images ---
- lscr.io/linuxserver/sonarr:mesh
- lscr.io/linuxserver/radarr:mesh
@@ -97,11 +158,11 @@ images:
- mesh-runtime-portainer:development
- mesh-runtime-minio:development
- mesh-runtime-mongodb:development
- mesh-runtime-lavinmq:development
- mesh-runtime-keycloak:development
- mesh-runtime-gitea:development
- mesh-runtime-nextcloud:development
- mesh-runtime-umami:development
- mesh-runtime-photos:development
- mesh-runtime-verdaccio:development
- mesh-runtime-mailu:development
- mesh-route-proxy:development