Rewrite the flat three-node whole-mesh-full (separate anchor, one public segment)
into production's real shape: two segments and one access point. novox sits on
the routable `hosting` segment and IS the anchor — it runs the substrate, its own
service set, the overlay hub and public ingress; there is no separate anchor node.
ace, shanks and g14 sit on the household `home` segment behind a NAT gateway,
reachable from outside only through what they dial out to.
The bed drives, and verifies, the thing the flat beds never could: the WireGuard
overlay forming ACROSS the access point — a home node dialling novox's public hub
endpoint out through the gateway's masquerade, the handshake completing through the
NAT, the keepalive holding the hole open. Phase A proves it (handshake state + a
ping over the overlay) before any heavy module lands; Phase B converges both server
sets. With MESH_LAB_KEEP the instance is raised under a fixed id and left standing.
Collapsing the substrate onto novox exposed real facts the separate-anchor beds
never hit, fixed here:
- the substrate bundle advertises the broker at 192.0.2.10 (the old anchor); a
token carries that verbatim as the endpoint a node dials, so with the substrate
on novox it must be novox's own public address. Rewritten at apply (the cert is
fingerprint-pinned, not hostname-checked, so only the address needs correcting).
- the two provider host-port collisions with the co-located substrate: postgres
5432 vs the store's 127.0.0.1:5432, lavinmq 5672 vs the broker's 127.0.0.1:5672.
Both provider host publishes are remapped off the substrate's ports.
And a lab limitation this first large-union bed exposed: the image registry VM took
the profile's default `dir` pool and a ~10GiB root, which the ~28GiB union of both
server sets overflows ("no space left on device"). raiseRegistry now places the
registry on the scenario's copy-on-write pool with a sized (default 80GiB, thin)
root disk, MESH_LAB_REGISTRY_DISK overridable.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Combine the novox (17-module) and ace (24-module) sets on ONE substrate and
prove both node-plans converge together. anchor runs the substrate only;
novox and ace each run their own self-contained set (own postgres/redis), so
nothing crosses a node boundary except enrolment and the shared broker/store.
The four modules both nodes run (postgres, redis, mssql, portainer) are added
once and assigned to each node, each getting its own per-node broker account.
An overlay is placed across all three nodes.
Proven green: both nodes converge together on the one substrate. ace reaches
applied+current with all 17 of its CORE up (and letta too this run); novox
reaches all 13 CORE up with its only failed resource the known firewall.load
oneshot gap. The two node-plans share one broker without collision — distinct
novox-<mod> and ace-<mod> accounts for the modules both run. No new cross-node
bug (overlay/DNS/identity/port) surfaced; ports are per-VM and the sets are
node-self-contained. Tolerates the same nine credential-sidecar gaps and
firewall's nftables.service oneshot documented in the per-server beds.
Resource envelope: 3 VMs (anchor 4GiB, novox 16GiB, ace 18GiB) + registry
scenery, ~79 union images (~35GB) stocked to one registry VM and pulled
concurrently by both nodes; fit within 125GiB host RAM and the 180GiB lab pool.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF