Commit Graph
7 Commits
Author SHA1 Message Date
jschoubben 5d6e8fbe7a Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben 555401a787 The bed bootstraps through the installer, not around it
ADR 0067's own acceptance check said the lab must raise its anchor by running the
program a bare machine runs. It did not: whole-mesh-full applied the substrate bundle
by hand and then looped enrolment over all four machines as one continuous operation.
That gets the order right by accident and models the wrong shape — and an install
procedure that exists only as a test fixture is exercised by whoever writes tests and
never by whoever installs, which is why every bootstrap fault this year was found late.

Two acts now, and the first gates the second.

  GENESIS is novox running mesh-bootstrap: the installer is built from source before
  the raise (make bootstrap, carrying the control-plane image built in the same run),
  placed beside the host binary, given the two manifests it reads, and run. The bed
  then asserts a WORKING MESH OF ONE — the control plane answers, the registry replies
  on /v2/, the container called mesh-control is running from a registry-pinned digest
  rather than an image id, the registry agrees it serves it, temp-mesh-control is gone,
  and the mesh has heard from its node. The image-id check is ADR 0067's "the pivot
  completed" verbatim: if it is still an id, nothing was published and this mesh can
  never roll out its own upgrades.

  JOINING is ace, shanks and g14: host binary, token, enrol, run. novox is NOT enrolled
  again — the installer already did it, and a second identity is one the mesh does not
  know.

If genesis stops, the bed prints which of the installer's ten steps it stopped at and
goes no further. A second machine joining a mesh that is not ready is a different
failure, and running it would bury this one underneath it.

The anchor is no longer handed mesh-control:development. Its absence is the point: the
installer carries that image inside itself, and handing it over as well would make the
load say "already held" and leave the carrying untested — the same class of fiction the
lab's own registry used to hide. A unit test asserts the scenario keeps it out.

The registry is reached at 127.0.0.1:5000, which is a finding rather than a shortcut: a
runtime refuses a plain-HTTP registry at any address but a loopback one, so the digest
the control-plane module is pinned to is one only the anchor can pull. Enough here,
because only the anchor runs a control plane. Written down in the bed.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 11:46:42 +02:00
jschoubben 94e617915c The home segment moves off 192.168.1.0/24
It is the commonest home LAN range there is, so on an ordinary workstation the
lab's private segment and the machine's own network are the same addresses. The
scenario routes an egress machine explicitly and marks the rest unreachable, so
nothing leaked — but that guard was carrying the whole weight of a collision
nobody chose, and a guard is a bad place for that.

10.99.1.0/24 is still RFC 1918, so the bed still models a home LAN behind an
access point. It is simply far from what this kind of machine already has:
192.168.1 is the LAN, 172.16-31 and 192.168.16-95 are container bridges, and
10.10/10.42/10.208 are a tunnel, the mesh overlay and the virtualisation daemon.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:00:19 +02:00
jschoubben 5c91c0ecd2 Scenarios: egress where images are needed, and images: cut to what is ours
Every scenario that places a container runtime gives each of its machines
`egress: true` — a node that runs modules pulls images from the internet, which is what
a node does. The underlay-only scenarios (bootstrap-single, behind-nat,
segmented-and-unforwardable, the-ordinary-shape, two-on-a-segment) stay sealed on
purpose: an extra NIC would change the very reachability they are asserting about.

`images:` keeps only the mesh's own — 55 third-party entries leave whole-mesh-full
alone, and the machine fetches them itself by the digest its module.json already pins.
The whole-mesh beds also say per machine which of ours they get: novox the substrate
control plane and its own fifteen runtimes, ace its twenty-four, the two workstations
one each. That is not a lab economy. An operator's workstation holds the images its own
modules need, and giving these two the union would put some thirty gigabytes onto a
thirty-gigabyte disk.

bootstrap-with-registry.yml is deleted. It existed only to demonstrate the lab's
registry, nothing referenced it, and there is nothing left for it to demonstrate.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:16:23 +02:00
jschoubben d9bf178546 whole-mesh-full: serve the internal CA's image
ADR 0056 made `acme-ca` a requirement of route-proxy, and step-ca is what
answers it. A scenario that does not stock the image cannot run the CA, the
proxy does not resolve, and every routed module on the mesh goes with it.

This was carried as an uncommitted edit through the first ADR 0056 raise. Kept,
because it is right, and committed, because a fix that lives in somebody's
working tree is a fix the next raise does not have.
2026-09-10 21:05:30 +02:00
jschoubben 80b0670ebe whole-mesh-full: the real segmented topology, and the overlay proven across the access point
Rewrite the flat three-node whole-mesh-full (separate anchor, one public segment)
into production's real shape: two segments and one access point. novox sits on
the routable `hosting` segment and IS the anchor — it runs the substrate, its own
service set, the overlay hub and public ingress; there is no separate anchor node.
ace, shanks and g14 sit on the household `home` segment behind a NAT gateway,
reachable from outside only through what they dial out to.

The bed drives, and verifies, the thing the flat beds never could: the WireGuard
overlay forming ACROSS the access point — a home node dialling novox's public hub
endpoint out through the gateway's masquerade, the handshake completing through the
NAT, the keepalive holding the hole open. Phase A proves it (handshake state + a
ping over the overlay) before any heavy module lands; Phase B converges both server
sets. With MESH_LAB_KEEP the instance is raised under a fixed id and left standing.

Collapsing the substrate onto novox exposed real facts the separate-anchor beds
never hit, fixed here:
- the substrate bundle advertises the broker at 192.0.2.10 (the old anchor); a
  token carries that verbatim as the endpoint a node dials, so with the substrate
  on novox it must be novox's own public address. Rewritten at apply (the cert is
  fingerprint-pinned, not hostname-checked, so only the address needs correcting).
- the two provider host-port collisions with the co-located substrate: postgres
  5432 vs the store's 127.0.0.1:5432, lavinmq 5672 vs the broker's 127.0.0.1:5672.
  Both provider host publishes are remapped off the substrate's ports.

And a lab limitation this first large-union bed exposed: the image registry VM took
the profile's default `dir` pool and a ~10GiB root, which the ~28GiB union of both
server sets overflows ("no space left on device"). raiseRegistry now places the
registry on the scenario's copy-on-write pool with a sized (default 80GiB, thin)
root disk, MESH_LAB_REGISTRY_DISK overridable.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 11:17:00 +02:00
jschoubben fcdcd338c0 Add whole-mesh full three-node bed (stage 3: both server sets, one substrate)
Combine the novox (17-module) and ace (24-module) sets on ONE substrate and
prove both node-plans converge together. anchor runs the substrate only;
novox and ace each run their own self-contained set (own postgres/redis), so
nothing crosses a node boundary except enrolment and the shared broker/store.
The four modules both nodes run (postgres, redis, mssql, portainer) are added
once and assigned to each node, each getting its own per-node broker account.
An overlay is placed across all three nodes.

Proven green: both nodes converge together on the one substrate. ace reaches
applied+current with all 17 of its CORE up (and letta too this run); novox
reaches all 13 CORE up with its only failed resource the known firewall.load
oneshot gap. The two node-plans share one broker without collision — distinct
novox-<mod> and ace-<mod> accounts for the modules both run. No new cross-node
bug (overlay/DNS/identity/port) surfaced; ports are per-VM and the sets are
node-self-contained. Tolerates the same nine credential-sidecar gaps and
firewall's nftables.service oneshot documented in the per-server beds.

Resource envelope: 3 VMs (anchor 4GiB, novox 16GiB, ace 18GiB) + registry
scenery, ~79 union images (~35GB) stocked to one registry VM and pulled
concurrently by both nodes; fit within 125GiB host RAM and the 180GiB lab pool.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-08 18:22:56 +02:00