Commit Graph
7 Commits
Author SHA1 Message Date
jschoubben 2f4cb871d9 A raise does not finish until the machines can pull from the registry
The first whole-mesh raise of the ADR 0056 code died on the anchor's substrate
apply: the image pulls failed, the anchor never came up, no node could enrol,
and the instance was left a bare shell — VMs and a registry, no substrate. The
identical apply, run by hand once the registry was warm, succeeded immediately.

`raiseRegistry` proves the wrong thing. It curls `localhost:5000` from inside
the registry's OWN machine, which says the registry process is up and holds the
blobs, and says nothing about the path anybody else uses: across a segment, and
for the home nodes through a NAT gateway whose default route and firewall are
applied two steps LATER. So "serving" was reported on evidence that excluded the
network, and the caller — which pins every image in the substrate bundle to that
registry — was handed a fact it could not rely on.

So the check moves to where it means something. After the routes and the
firewalls, before the minutes spent placing, each machine is asked for `/v2/` and
for one stocked manifest BY DIGEST, at the address it will pin, over the network
it will use. That is the pair of requests a pull begins with, from the same
place. Layers are not fetched: every digest was already read back inside the
registry machine, so what is in question here is the path, not the content.

Verified by typecheck and the unit suite (136 pass), and by confirming against a
standing four-node instance that `curl` exists in the machines and that both
segments — including a home node through the gateway — answer 200 for the
registry's `/v2/`. The ordering itself is unverified in a live raise from cold,
which takes hours.
2026-09-10 21:05:52 +02:00
jschoubben 80b0670ebe whole-mesh-full: the real segmented topology, and the overlay proven across the access point
Rewrite the flat three-node whole-mesh-full (separate anchor, one public segment)
into production's real shape: two segments and one access point. novox sits on
the routable `hosting` segment and IS the anchor — it runs the substrate, its own
service set, the overlay hub and public ingress; there is no separate anchor node.
ace, shanks and g14 sit on the household `home` segment behind a NAT gateway,
reachable from outside only through what they dial out to.

The bed drives, and verifies, the thing the flat beds never could: the WireGuard
overlay forming ACROSS the access point — a home node dialling novox's public hub
endpoint out through the gateway's masquerade, the handshake completing through the
NAT, the keepalive holding the hole open. Phase A proves it (handshake state + a
ping over the overlay) before any heavy module lands; Phase B converges both server
sets. With MESH_LAB_KEEP the instance is raised under a fixed id and left standing.

Collapsing the substrate onto novox exposed real facts the separate-anchor beds
never hit, fixed here:
- the substrate bundle advertises the broker at 192.0.2.10 (the old anchor); a
  token carries that verbatim as the endpoint a node dials, so with the substrate
  on novox it must be novox's own public address. Rewritten at apply (the cert is
  fingerprint-pinned, not hostname-checked, so only the address needs correcting).
- the two provider host-port collisions with the co-located substrate: postgres
  5432 vs the store's 127.0.0.1:5432, lavinmq 5672 vs the broker's 127.0.0.1:5672.
  Both provider host publishes are remapped off the substrate's ports.

And a lab limitation this first large-union bed exposed: the image registry VM took
the profile's default `dir` pool and a ~10GiB root, which the ~28GiB union of both
server sets overflows ("no space left on device"). raiseRegistry now places the
registry on the scenario's copy-on-write pool with a sized (default 80GiB, thin)
root disk, MESH_LAB_REGISTRY_DISK overridable.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 11:17:00 +02:00
jschoubben bb14ecb7e0 Say what the lab is doing, while it is doing it
novox/hq 04-ISSUES/024. A run stalled for thirty-five minutes and said
nothing. The cause was a link systemd was still configuring, three
layers down inside a `docker load` blocked on a socket — and every one
of those layers knew what it was waiting for. None of them said so.

Three decisions, each doing work.

**Every external command is logged, at the three places that run one.**
Ninety-seven call sites reach a hypervisor or a container runtime
through three wrappers, so instrumenting the wrappers covers all of them
and nothing has to remember to log.

**A command still running says so while it runs.** A line before and a
line after tells you nothing until the after arrives, which is exactly
the case that matters. Anything outstanding past fifteen seconds reports
itself with how long it has been going. It is reported as still running,
not as stuck — which it is is not knowable from there, and a log that
calls a slow step a hang teaches people to ignore it.

**It goes to a file, written synchronously.** Node block-buffers stdout
when redirected and a test runner buffers it again, so a console log can
sit minutes behind. `appendFileSync` cannot lag.

Two things this found in itself while being written, both the same shape
as what it exists to catch:

A question that answers no is not a fault. Half the lab's commands are
questions — does this network exist, is the agent up yet — and they fail
constantly while a scenario comes up. Logging those as faults filled a
healthy run with ✗, which is how you end up ignoring ✗ when one is real.
They are recorded quietly now, and still recorded.

And `around` skipped its own wrapper when a step's level was below the
configured one — taking the failure line and the heartbeat with it. The
two things worth having at a low level were the two that vanished at
exactly the level somebody would use. The gate belongs in `write`.

Also unsilences the four call sites that passed a callback throwing
everything away, including the one the stall sat in, and tees `raise`'s
progress into the file whether or not a caller asked to see it — the
end-to-end test passed no callback, so the one run that mattered
reported not a single step.
2026-09-01 10:43:38 +02:00
jschoubben 3503ad990b The registry is addressed the way every other machine is
novox/hq 04-ISSUES/024. The registry machine had its address set with
`ip addr add`; every other machine gets a systemd-networkd unit. That
one difference stalled the lab indefinitely.

An address set by hand leaves networkd waiting to configure a link it
was never told about, so the link sits at `configuring` for ever.
`systemd-networkd-wait-online` has TimeoutStartUSec=infinity, so
`network-online.target` is never reached — and Docker is ordered after
it. `docker load` then blocked on a socket whose daemon was queued
behind a target that would never come.

Measured before and after on the same scenario: stuck with five pending
systemd jobs and `docker` inactive; now `enp5s0 configured`, `docker`
active, no jobs, and the whole raise completes in 87.5s.

The guess in the issue was wrong, and it was wrong in the usual way —
stocking had just been changed, so stocking looked guilty. Stocking
takes 34s and always did.

Two things that made this cost hours rather than minutes are fixed with
it. Placing an image now waits for the container runtime to answer and
refuses after 120s naming what systemd is waiting on, so a stall becomes
a failure that says why instead of three stacked timeouts totalling 35
minutes. And the end-to-end test passes `onProgress`, so a raise says
what step it is on — it printed nothing at all until it finished, which
is why 35 minutes of nothing read as a slow test.
2026-09-01 09:37:16 +02:00
jschoubben be176bab2e Automate the lab registry: a sealed machine pulls by digest
Closes 04-ISSUES/009. A scenario declares `images:` by tag; the lab stocks a
registry on this workstation where there is a network, raises it inside the
scenario as scenery, and reports the references a declaration pins -- which are
the digests THIS registry assigned, and are not knowable until it is raised.

Verified in a sealed machine, confirmed by ping to have no route out: package,
service including boot state, a container pinned by digest, and an action
inside that container. Applied, idempotent on re-apply, and read back from the
machine rather than from the apply's own report. That is the first time the
container shape has worked in the lab at all, and it was the shape blocking the
substrate bootstrap.

Four faults found by running it, three of them mine and one worth keeping:

The read-back checked that the catalog endpoint answered, by looking for the
substring "repositories" -- which `{"repositories":[]}` also contains. So it
passed on a registry holding nothing, and the failure surfaced much later as a
container that could not be pulled. It now asks for each image's manifest BY
DIGEST, which is what a machine does.

A recursive push needs its destination to exist, or incus copies the source's
contents rather than the source. The data landed one directory too shallow and
the registry found nothing where it looks.

The registry writes its blobs as root through a bind mount, so the workstation
could not remove its own scratch directory afterwards. Whoever made the files
removes them -- the cleanup now runs in a container too. And a cleanup failure
no longer fails a raise that succeeded: the scenario is standing and usable,
and saying otherwise would be a false report.

The base image build did not verify that the runtime trusts the documentation
ranges as plain-HTTP registries. Writing the file is not the daemon honouring
it, and a base image that looks right fails much later, in a sealed scenario,
a long way from its cause. It is now read back from `docker info`.
2026-08-29 00:04:55 +02:00
jschoubben 4097ff92c1 Repoint ADR references after HQ consolidated 65 records to 23
Comments naming records that no longer exist now point at the consolidated
record holding their reasoning -- the four lab records are 0016, a test defends
a decision is 0017.
2026-08-28 23:33:46 +02:00
jschoubben 37c6a0ba21 A registry inside the scenario: the mechanism, verified
Issue 009's resolution, proven manually end to end before any of it was
written.

A sealed machine pulled an image BY DIGEST from a registry on its own segment
and ran it; then the host applied all four shapes -- package, service with
boot, container from that digest, and an action inside it -- idempotently. That
is the first time the container shape has worked anywhere but a workstation,
and it was the shape blocking the whole substrate bootstrap.

The registry's digests are its own, not Docker Hub's, and that is correct
rather than a compromise: ADR 0046 requires a reference that is exact and
cannot move, and a digest this registry assigned is both. It is also not a
lab workaround -- 0048 names an OCI registry as substrate and 0046 says a first
node fetches "upstream, wherever the image ordinarily lives". This IS that
upstream, scenery in the same sense the transit router is the internet.

The base image now trusts the RFC 5737 and RFC 3849 documentation ranges as
plain-HTTP registries. Scoped to those rather than an address because they
never route on the real internet, so it cannot make a real machine trust a real
registry whatever it is copied onto.

Three faults found while verifying, two of them mine:

My probe script picked an interface with `ls /sys/class/net | head -1`, which
returns docker0 once a runtime exists -- so it addressed the wrong interface and
then, because that address overlapped the segment, broke routing on the machine
entirely. The lab itself is immune: it matches by MAC, for a related reason it
already recorded (bus-position naming on multi-homed machines).

And a test that proved nothing: I asserted `sha256:tooshort` is rejected, but
its letters fall outside a-f, so it failed the character class rather than the
length check. Replaced with hex of the wrong length, after which removing the
length check bites.
2026-08-28 02:18:35 +02:00