whole-mesh-full: the real segmented topology, and the overlay proven across the access point
Rewrite the flat three-node whole-mesh-full (separate anchor, one public segment)
into production's real shape: two segments and one access point. novox sits on
the routable `hosting` segment and IS the anchor — it runs the substrate, its own
service set, the overlay hub and public ingress; there is no separate anchor node.
ace, shanks and g14 sit on the household `home` segment behind a NAT gateway,
reachable from outside only through what they dial out to.
The bed drives, and verifies, the thing the flat beds never could: the WireGuard
overlay forming ACROSS the access point — a home node dialling novox's public hub
endpoint out through the gateway's masquerade, the handshake completing through the
NAT, the keepalive holding the hole open. Phase A proves it (handshake state + a
ping over the overlay) before any heavy module lands; Phase B converges both server
sets. With MESH_LAB_KEEP the instance is raised under a fixed id and left standing.
Collapsing the substrate onto novox exposed real facts the separate-anchor beds
never hit, fixed here:
- the substrate bundle advertises the broker at 192.0.2.10 (the old anchor); a
token carries that verbatim as the endpoint a node dials, so with the substrate
on novox it must be novox's own public address. Rewritten at apply (the cert is
fingerprint-pinned, not hostname-checked, so only the address needs correcting).
- the two provider host-port collisions with the co-located substrate: postgres
5432 vs the store's 127.0.0.1:5432, lavinmq 5672 vs the broker's 127.0.0.1:5672.
Both provider host publishes are remapped off the substrate's ports.
And a lab limitation this first large-union bed exposed: the image registry VM took
the profile's default `dir` pool and a ~10GiB root, which the ~28GiB union of both
server sets overflows ("no space left on device"). raiseRegistry now places the
registry on the scenario's copy-on-write pool with a sized (default 80GiB, thin)
root disk, MESH_LAB_REGISTRY_DISK overridable.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
This commit is contained in:
@@ -321,7 +321,7 @@ export async function raise(
|
||||
let registry: Awaited<ReturnType<typeof raiseRegistry>> = null;
|
||||
try {
|
||||
enter("raising the registry");
|
||||
registry = await raiseRegistry(scenario, instanceId, stock, log);
|
||||
registry = await raiseRegistry(scenario, instanceId, stock, log, pool);
|
||||
} finally {
|
||||
// Cleaning up scratch must not fail a raise that succeeded. The scenario is standing
|
||||
// and usable; a directory left behind is untidy, and saying so is the honest report.
|
||||
|
||||
@@ -276,6 +276,13 @@ export async function raiseRegistry(
|
||||
instanceId: string,
|
||||
stock: Stock,
|
||||
log: (message: string) => void = () => {},
|
||||
/**
|
||||
* The copy-on-write pool the scenario's machines were placed on. The registry goes on it too —
|
||||
* NOT the profile's default `dir` pool — so a sized root disk is thin (paid for as it fills)
|
||||
* rather than a full allocation on the host's own filesystem. Absent, it falls back to the
|
||||
* profile default, which is the historical behaviour for a small registry.
|
||||
*/
|
||||
pool?: string,
|
||||
): Promise<RaisedRegistry | null> {
|
||||
if (stock.images.length === 0) return null;
|
||||
|
||||
@@ -292,10 +299,19 @@ export async function raiseRegistry(
|
||||
const prefix = segment.cidr.slice(segment.cidr.lastIndexOf("/"));
|
||||
|
||||
if (!(await succeeds(["config", "show", name], 15_000))) {
|
||||
// The registry holds the WHOLE `images:` union on its own root disk, and the base image's
|
||||
// default is only ~10GiB. A single-node bed stocks a handful of images and fits; a broad bed
|
||||
// — and especially the full segmented mesh, whose union is both server sets at once (~28GiB)
|
||||
// — overflows it, and the raise dies "no space left on device" while pushing blobs into the
|
||||
// registry. So the registry gets a sized root disk. Thin on a copy-on-write pool, so a small
|
||||
// bed pays only for what it actually stocks; MESH_LAB_REGISTRY_DISK overrides for a giant one.
|
||||
const registryDisk = process.env["MESH_LAB_REGISTRY_DISK"] ?? "80GiB";
|
||||
await incus([
|
||||
"init", BASE_IMAGE_ALIAS, name, "--vm",
|
||||
"-c", "security.secureboot=false",
|
||||
"-c", "limits.memory=1GiB",
|
||||
...(pool ? ["-s", pool] : []),
|
||||
"-d", `root,size=${registryDisk}`,
|
||||
"-c", `user.mesh-lab.instance=${instanceId}`,
|
||||
// Tagged as a machine as well, so `destroy` finds it with one query — a router that
|
||||
// carried only its own tag was left behind and held its networks open.
|
||||
|
||||
Reference in New Issue
Block a user