Commit Graph
14 Commits
Author SHA1 Message Date
jschoubben fb6dab6a48 The beds place the builder's manifest too 2026-09-14 12:31:44 +02:00
jschoubben d27f24cf3e The bed resolves an artifact the way the builder would
It pre-builds these images and stocks them, which is the lab standing in for the
builder — so it must do what the builder does and replace the artifact with the
reference the machine holds. Without it the unresolved field travels to the
machine and the whole declaration is refused.
2026-09-13 04:56:00 +02:00
jschoubben fb18807000 The bed checks the control plane was built, not carried
The pivot checks proved the running control plane is pinned to a digest this
mesh's registry serves, which a carried image satisfies just as well. What the
installer now exists to make true is that a build happened, from the commit the
bed asked for — and that was printed and not checked.
2026-09-13 04:28:54 +02:00
jschoubben 607ea241c7 The lab's installer carries a builder, and genesis is told what to build
Both beds now pass a repository and a commit, and refuse to run without them
rather than raising a machine the installer cannot finish.
2026-09-13 04:24:18 +02:00
jschoubben 3f58a0a08f The provider remap moves a port; it does not bind it to loopback
postgres and lavinmq carried 127.0.0.1: in their remap, and it broke a consumer
on a node that has no substrate to collide with. A module is told to reach its
provider at <node>.internal, that name is the node's overlay address, and a
provider listening only on loopback refuses it — letta on ace failed with 'is the
server running on that host and accepting TCP/IP connections?' while postgres sat
healthy beside it.

The collision needed a different port, which is what every other entry here does.
The address was never part of it, and it made the provider unreachable by the one
name the mesh hands its consumers.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 22:05:18 +02:00
jschoubben 555401a787 The bed bootstraps through the installer, not around it
ADR 0067's own acceptance check said the lab must raise its anchor by running the
program a bare machine runs. It did not: whole-mesh-full applied the substrate bundle
by hand and then looped enrolment over all four machines as one continuous operation.
That gets the order right by accident and models the wrong shape — and an install
procedure that exists only as a test fixture is exercised by whoever writes tests and
never by whoever installs, which is why every bootstrap fault this year was found late.

Two acts now, and the first gates the second.

  GENESIS is novox running mesh-bootstrap: the installer is built from source before
  the raise (make bootstrap, carrying the control-plane image built in the same run),
  placed beside the host binary, given the two manifests it reads, and run. The bed
  then asserts a WORKING MESH OF ONE — the control plane answers, the registry replies
  on /v2/, the container called mesh-control is running from a registry-pinned digest
  rather than an image id, the registry agrees it serves it, temp-mesh-control is gone,
  and the mesh has heard from its node. The image-id check is ADR 0067's "the pivot
  completed" verbatim: if it is still an id, nothing was published and this mesh can
  never roll out its own upgrades.

  JOINING is ace, shanks and g14: host binary, token, enrol, run. novox is NOT enrolled
  again — the installer already did it, and a second identity is one the mesh does not
  know.

If genesis stops, the bed prints which of the installer's ten steps it stopped at and
goes no further. A second machine joining a mesh that is not ready is a different
failure, and running it would bury this one underneath it.

The anchor is no longer handed mesh-control:development. Its absence is the point: the
installer carries that image inside itself, and handing it over as well would make the
load say "already held" and leave the carrying untested — the same class of fiction the
lab's own registry used to hide. A unit test asserts the scenario keeps it out.

The registry is reached at 127.0.0.1:5000, which is a finding rather than a shortcut: a
runtime refuses a plain-HTTP registry at any address but a loopback one, so the digest
the control-plane module is pinned to is one only the anchor can pull. Enough here,
because only the anchor runs a control plane. Written down in the bed.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 11:46:42 +02:00
jschoubben 94e617915c The home segment moves off 192.168.1.0/24
It is the commonest home LAN range there is, so on an ordinary workstation the
lab's private segment and the machine's own network are the same addresses. The
scenario routes an egress machine explicitly and marks the rest unreachable, so
nothing leaked — but that guard was carrying the whole weight of a collision
nobody chose, and a guard is a bad place for that.

10.99.1.0/24 is still RFC 1918, so the bed still models a home LAN behind an
access point. It is simply far from what this kind of machine already has:
192.168.1 is the LAN, 172.16-31 and 192.168.16-95 are container bridges, and
10.10/10.42/10.208 are a tunnel, the mesh overlay and the virtualisation daemon.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:00:19 +02:00
jschoubben a4c2a9b90b The routing record is 0066, not 0056
0056 was already 'the authority is the control plane, not a database'. The
routing record was renumbered where it lives; these citations pointed at the
wrong decision, which is worse than pointing at none.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:35:35 +02:00
jschoubben 675facdb0d The beds name images the way a machine would find them
Twenty-eight integration tests each carried their own copy of the same two helpers,
which pointed a manifest and the substrate bundle at whatever the lab's registry had
assigned. They now share two in the harness, and the difference is the point: ours is
rewritten to the ID the machine holds it under, and everything else is left exactly as
written so the machine pulls it.

**The substrate bundle is where the fiction was most load-bearing.** mesh-host's
`examples/substrate-first-node.lock` pins all three of its images at
`192.0.2.250:5000/…`, which is the address the lab's registry served from — it was
written for a target, and the target was the lab. Two of those are ordinary third-party
images and become the digests mesh-catalog's own postgres and lavinmq modules pin, so
the substrate's store and broker are literally the images the mesh runs. mesh-control
exists in no registry at all and becomes the ID the machine was handed. **The bundle
itself should be fixed in mesh-host and this substitution deleted with it.**

Beds that wrote a manifest by hand named an image by repository and let the rewrite
supply a digest. There is nothing to supply one now, so `onTheMachine` refuses an
unpinned reference and hands back the digest the catalogue pins — a bed runs the image
the mesh ships, and a bed that drifts from the catalogue is testing a different
postgres.

Three beds took a third-party image out of the raised list, which no longer contains
one: certificates (pebble), objectstore (minio and its client) and provisioner
(postgres) now name theirs and pull it. builds and mesh publish into the MESH's own
artifact store — the `registry` module's image, on the node, on 5000 — rather than into
scenery the lab raised. That is a different claim, and only one of them exists in
production.

New unit tests cover what a full raise would otherwise be the only way to check: the
routes an egress machine gets (that its gateway is still the path to the rest of the
scenario, that a range with no path is unreachable rather than leaked to the uplink,
that each family gets its own next hop), which machine is handed which of our images,
and the `images:` rule that refuses a third-party entry. The "shipped scenarios are
valid" test now loads every scenario rather than two of them.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:16:41 +02:00
jschoubben 0a0c57b610 whole-mesh-full: the CA root a person hands the mesh must be readable by it
The control plane's image is FROM scratch and runs as 65534, and docker cp keeps
the mode a file had outside — openssl writes a private key 0600 root-owned, so
the copy landed unreadable, secret accept failed with permission denied, and the
CA crash-looped on a root it never got. Chowning it inside the container is not
available: there is no shell in there to do it with.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 22:35:49 +02:00
jschoubben a48b60b604 whole-mesh-full: the bed knows about ADR 0056
The bed set no node a `public-domain` and assigned no `acme-ca` provider, so it
was testing a mesh the design no longer describes — and going green while doing
it, which is the worse half.

**No public domain means no route.** A module now contributes a `label` and
nothing else; the mesh joins it to the node's public domain, and a label with no
domain to join composes to nothing at all. Every routed module on this bed was
therefore unreachable by name, silently, and no assertion noticed. novox now
carries `novox.incus` and ace `zurag.incus` — `.incus`, because this repository's
beds name nothing routable. The workstations carry none, which is also the design
being exercised: a node that does not face outward has no public domain.

**No acme-ca provider means no proxy.** route-proxy requires one, so without a
provider it is unresolvable and takes every routed module with it. step-ca is
assigned on the anchor, at mesh scope, and given an operator root — made with
openssl on the anchor and handed over through the real `secret accept` path,
because the mesh cannot invent a PEM and the random bytes it makes for an
own-secret nobody supplied would leave the CA crash-looping on a root key that is
not a key.

**What is asserted is the half that is decided and cheap**: that each routed
module's name composes to `<label>.<public-domain>` — read from the proxy's own
received-routes file, the mesh's answer on the machine rather than this test's
arithmetic checked against itself — with `@` composing to the bare domain, and
that the proxy answers for one of them over HTTP.

**What is NOT asserted is issuance.** Whether route-proxy obtains a certificate
from step-ca over ACME depends on mesh-control fixes landing as this is written,
and a bed that gated on them would report somebody else's in-flight work as its
own failure. step-ca is listed as a reported gap for the same reason.

The substrate apply also retries up to three times. `raise` now refuses to return
until every machine can fetch a manifest from the scenario registry, so the first
attempt should be the only one; a pull is simply the one step here that can fail
for a reason that goes away by itself, and the cost of not retrying was a whole
raise left as a bare shell.

Typechecks; not run end-to-end — see the ADR 0056 section for what is expected to
fail until the issuance path is fixed.
2026-09-10 21:06:34 +02:00
jschoubben 80b0670ebe whole-mesh-full: the real segmented topology, and the overlay proven across the access point
Rewrite the flat three-node whole-mesh-full (separate anchor, one public segment)
into production's real shape: two segments and one access point. novox sits on
the routable `hosting` segment and IS the anchor — it runs the substrate, its own
service set, the overlay hub and public ingress; there is no separate anchor node.
ace, shanks and g14 sit on the household `home` segment behind a NAT gateway,
reachable from outside only through what they dial out to.

The bed drives, and verifies, the thing the flat beds never could: the WireGuard
overlay forming ACROSS the access point — a home node dialling novox's public hub
endpoint out through the gateway's masquerade, the handshake completing through the
NAT, the keepalive holding the hole open. Phase A proves it (handshake state + a
ping over the overlay) before any heavy module lands; Phase B converges both server
sets. With MESH_LAB_KEEP the instance is raised under a fixed id and left standing.

Collapsing the substrate onto novox exposed real facts the separate-anchor beds
never hit, fixed here:
- the substrate bundle advertises the broker at 192.0.2.10 (the old anchor); a
  token carries that verbatim as the endpoint a node dials, so with the substrate
  on novox it must be novox's own public address. Rewritten at apply (the cert is
  fingerprint-pinned, not hostname-checked, so only the address needs correcting).
- the two provider host-port collisions with the co-located substrate: postgres
  5432 vs the store's 127.0.0.1:5432, lavinmq 5672 vs the broker's 127.0.0.1:5672.
  Both provider host publishes are remapped off the substrate's ports.

And a lab limitation this first large-union bed exposed: the image registry VM took
the profile's default `dir` pool and a ~10GiB root, which the ~28GiB union of both
server sets overflows ("no space left on device"). raiseRegistry now places the
registry on the scenario's copy-on-write pool with a sized (default 80GiB, thin)
root disk, MESH_LAB_REGISTRY_DISK overridable.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 11:17:00 +02:00
jschoubben 591f2a641c whole-mesh-full: prove the dry-run fixes (fail2ban hostable, credential own-secrets)
Re-runs the capstone from main after the dry-run fixes merged.

fail2ban: added to the novox set. The capability fix (intrusion-prevention ->
firewall) makes it HOSTABLE — it is now assigned, not refused — which is the
gate. Its service reaching active is a host concern the offline lab cannot meet
(the VM ships nftables but not fail2ban, and the isolated segment has no route to
the package mirror, so pacman cannot fetch it), so fail2ban joins GAPS_NOVOX: its
failed package resource is tolerated like firewall's oneshot nftables.service.

7 credential sidecars: before the push, a FAKE app credential is delivered for
each (plex/bazarr/ombi/home-assistant/nzbget/qbittorrent on ace, umami on novox)
through the real operator path — `secret accept <node> <module> <name> --from`.
The bed asserts each sidecar advances PAST its old "no credential" crash (it reads
the delivered value); app-auth failure against the real app with a bogus value is
expected and not gated.

Result: SUITE_EXIT=0. Both node-plans converge on one substrate (novox 13/13
core, ace 17/17 core), fail2ban hostable, all 7 sidecars past their crash.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-08 20:05:04 +02:00
jschoubben fcdcd338c0 Add whole-mesh full three-node bed (stage 3: both server sets, one substrate)
Combine the novox (17-module) and ace (24-module) sets on ONE substrate and
prove both node-plans converge together. anchor runs the substrate only;
novox and ace each run their own self-contained set (own postgres/redis), so
nothing crosses a node boundary except enrolment and the shared broker/store.
The four modules both nodes run (postgres, redis, mssql, portainer) are added
once and assigned to each node, each getting its own per-node broker account.
An overlay is placed across all three nodes.

Proven green: both nodes converge together on the one substrate. ace reaches
applied+current with all 17 of its CORE up (and letta too this run); novox
reaches all 13 CORE up with its only failed resource the known firewall.load
oneshot gap. The two node-plans share one broker without collision — distinct
novox-<mod> and ace-<mod> accounts for the modules both run. No new cross-node
bug (overlay/DNS/identity/port) surfaced; ports are per-VM and the sets are
node-self-contained. Tolerates the same nine credential-sidecar gaps and
firewall's nftables.service oneshot documented in the per-server beds.

Resource envelope: 3 VMs (anchor 4GiB, novox 16GiB, ace 18GiB) + registry
scenery, ~79 union images (~35GB) stocked to one registry VM and pulled
concurrently by both nodes; fit within 125GiB host RAM and the 180GiB lab pool.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-08 18:22:56 +02:00