Deleting the lab's registry left the operator's own images to be pulled like
anything else, and they cannot be: their registry wants an account and a
scenario machine has none. The pull fails with 'no basic auth credentials',
which is not something more patience fixes.
So the test is no longer 'did the mesh build it' but 'can the machine get it at
all'. Two ways to fail that — published nowhere, or published somewhere the
machine cannot authenticate to — and one consequence: the workstation, which
does hold the credential, exports it and loads it.
Worth saying what this stands in for. In a finished mesh these are built by the
builder and published to the mesh's own store, and every machine pulls them from
there with a credential the mesh granted. Until that store exists there is
nowhere for them to come from, and handing them over is the closest honest thing
— not a registry the lab invents, which is what was just removed.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Seven catalogue modules — photos, photos-eef, photos-filip, invoicing, novox.be,
de-spiegel, amqp-email-forwarder — name `registry-api.…/novox/…:latest`. That is a TAG,
which ADR 0006 forbids and mesh-host refuses. It has never shown, because the lab's
registry rewrote every reference to a digest it had assigned, tag or not: the fiction
was not only serving the images, it was silently pinning them.
There is nothing to pin them with now. Asserting here would take whole-mesh-full down in
`before()`, before the overlay it exists to prove; the useful outcome is that each of
those modules fails to apply on the node that carries it, saying exactly why, while the
rest of the bed runs. So the reference passes through and the harness says so out loud.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Twenty-eight integration tests each carried their own copy of the same two helpers,
which pointed a manifest and the substrate bundle at whatever the lab's registry had
assigned. They now share two in the harness, and the difference is the point: ours is
rewritten to the ID the machine holds it under, and everything else is left exactly as
written so the machine pulls it.
**The substrate bundle is where the fiction was most load-bearing.** mesh-host's
`examples/substrate-first-node.lock` pins all three of its images at
`192.0.2.250:5000/…`, which is the address the lab's registry served from — it was
written for a target, and the target was the lab. Two of those are ordinary third-party
images and become the digests mesh-catalog's own postgres and lavinmq modules pin, so
the substrate's store and broker are literally the images the mesh runs. mesh-control
exists in no registry at all and becomes the ID the machine was handed. **The bundle
itself should be fixed in mesh-host and this substitution deleted with it.**
Beds that wrote a manifest by hand named an image by repository and let the rewrite
supply a digest. There is nothing to supply one now, so `onTheMachine` refuses an
unpinned reference and hands back the digest the catalogue pins — a bed runs the image
the mesh ships, and a bed that drifts from the catalogue is testing a different
postgres.
Three beds took a third-party image out of the raised list, which no longer contains
one: certificates (pebble), objectstore (minio and its client) and provisioner
(postgres) now name theirs and pull it. builds and mesh publish into the MESH's own
artifact store — the `registry` module's image, on the node, on 5000 — rather than into
scenery the lab raised. That is a different claim, and only one of them exists in
production.
New unit tests cover what a full raise would otherwise be the only way to check: the
routes an egress machine gets (that its gateway is still the path to the rest of the
scenario, that a range with no path is unreachable rather than leaked to the uplink,
that each family gets its own next hop), which machine is handed which of our images,
and the `images:` rule that refuses a third-party entry. The "shipped scenarios are
valid" test now loads every scenario rather than two of them.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Comments naming records that no longer exist now point at the consolidated
record holding their reasoning -- the four lab records are 0016, a test defends
a decision is 0017.
The suite next door raises one scenario and asks deep questions of it. This one
asks shallow questions of every scenario — the half that was missing, since both
faults found by hand lived in scenarios nothing ever built.
Adds bootstrap-single, the cheapest, and the loop that lets the list grow. Also
adds the second universal invariant: every address a scenario declared is one
the machine actually holds. A machine that came up bare looks identical to one
that came up correctly until something asks it.
Verified to bite rather than assumed: against a live instance, the real
declaration passes and a declaration claiming an address nothing holds fails
with 'anchor declared 192.0.2.99 on hosting but holds 192.0.2.10'.
Integration now runs with --test-concurrency=1. Two files raise real instances,
node --test runs files in parallel by default, and two concurrent runs of this
suite already produced a whole-suite failure once — every test red, from
resource contention rather than from any fault in the code.
Gate: 45.7s -> 60.2s.
Reviewed and the criticism was right: 1,072 of 2,128 lines untested, all of
it the half that touches the hypervisor, and no gate. The verification I had
done was real — pings across NAT, TTL counts, ruleset comparisons — and none
of it survived the terminal it ran in, which is 04-ISSUES/005 in miniature.
Ten integration tests against a real hypervisor, each named for what it
defends. ADR 0031: a raised machine carries no overlay, no wireguard, no
mesh config — a scenario that pre-built peering would certify its own work.
ADR 0032: exec is the only way in. ADR 0033: routers are containers while
machines are virtual machines. And the design's claims: raise waits for
usable, snapshots are whole-scenario, NAT hides a private address,
published reaches the machine at the gateway's address.
Mocking the hypervisor is forbidden, so they skip with a reason on a
machine that cannot raise scenarios rather than passing green having
checked nothing.
The suite earned itself on its first run. It found that a snapshot of a
running machine could miss a file written seconds earlier — not stale,
absent — because the write was still in the guest's page cache. That is
exactly the question the lifecycle design listed as open: does a scenario
snapshot need the machines stopped? It does not, but it does need them
flushed. snapshot now syncs every machine before capturing, and the design
records the answer.
The fix buys write-durability, not application-consistency: anything
mid-transaction is still captured mid-transaction, and that is now stated
rather than assumed.
npm run check is the gate — typecheck, 40 unit tests, 10 integration tests.