The lab raised a `registry` VM, pushed ~73 images into it from the workstation, and
rewrote every manifest reference — third-party ones included — to point at it. No
production mesh has such a thing. So every bed proved that a machine could fetch an
image from a registry that exists nowhere else, and the bootstrap problems that only
appear when a machine has to fetch for itself went unfound.
What replaces it is the two things that are true in the world:
**Public images come from the public internet.** mesh-lab already created a NAT'd
uplink for exactly this and attached it to any machine declaring `egress`; no scenario
ever declared it. They do now, and third-party references are left exactly as the
catalogue writes them.
**The mesh's own images have no registry and never will.** mesh-control, mesh-builder,
mesh-route-proxy and the per-module runtimes are built from source and exist in no
registry. A machine gets them the way an operator's machine does — they are built here
and loaded onto it — and is then named by the digest of its own image configuration,
which mesh-host now accepts as "an image this machine already holds".
`images:` therefore means only *ours*, and a third-party entry is refused rather than
quietly loaded: otherwise the fiction returns one convenient line at a time. It is
per-machine as well, because "everything, everywhere" was never a description of
anything real — handing whole-mesh-full's union to its two 30GiB workstations would
fill the disk with runtimes nothing on them will start.
**The uplink and the declared gateway would have fought, silently.** A gateway container
and the transit router reach the scenario and nothing else; a default route through
either is a black hole for anything outside, and it beats the uplink's DHCP route on
metric. So a machine with egress states the scenario's ranges explicitly — through the
same gateway or transit it would have defaulted to, so the overlay-across-NAT path is
unchanged — and leaves the default to the uplink. A range with no path inside the
scenario becomes `unreachable` rather than falling through: 192.168.1.0/24 is an
ordinary private range in fact, and letting it escape would put scenario traffic on
whatever network the workstation is sitting on. `scenarioRoutesFor` is pure and tested,
because a decision only a full raise could check is one nobody checks.
The registry-reachability check the raise gained earlier is kept, pointed at the real
thing: every machine with egress must resolve a name and reach the internet before the
raise says it finished. Same failure it was written for — a raise that returns, an apply
that dies on its first pull, an instance left a bare shell — now guarding the path that
actually carries.
The base image's trust of the documentation ranges as plain-HTTP registries STAYS. It
was never only for the lab's registry: the mesh has one of its own, the `registry`
module, serving artifacts to the whole mesh over plain HTTP from whatever node runs it.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
whatWasTested read the repositories when the run ended, so a commit
landing during the twenty minutes a suite takes was recorded as tested
without ever being in the binaries. It happened: one receipt named a
commit made mid-run, and the verdict it carried belonged to an older
tree.
The heads are read once, right after the build, and carried to the
receipt. A verdict is only worth something attributed to one exact
state, which is the receipt's whole reason to exist.
The canary walked one path on one machine — a mesh comes up, a module
lands, a consumer gets a credential — and stopped the run if it broke.
That path is exactly what the first three tests of the long run walk,
and the long run finishes them about 160 seconds in.
So the gate cost a whole scenario on every passing run to save roughly
45 seconds on a failing one. A scenario is three machines, one of them a
registry that boots a kernel in order to serve files, which is where the
two minutes went.
The test file stays and still runs when it is named. What is gone is
raising it on the way to everything else.
Measured rather than argued: the canary's scenario took 116s of which
60s was standing up a registry, and the run reached the same assertions
without it.
Suggested by Jochen, and it paid for itself on its first run.
A suite that takes forty minutes is a suite you hear from once a day.
Every fault found today would have shown up in the first three minutes
of it — a module pinned to an image that does not exist, a consumer
given a password and no name to present with it, a credential file
nothing could read, a search for a password that read the password as an
option. The other thirty-seven minutes proved things that were already
working.
So this runs first, on one machine, with the three images the mesh needs
for itself. It walks one path: a mesh comes up, a module lands, and a
consumer gets a credential it can actually use — the name to present,
the address, the port, and a password only the host could put there.
Deliberately not a smaller copy of the full suite: that path is where
everything went wrong, and a canary checking many things shallowly is a
canary whose failure nobody can read.
`suite` runs it and stops if it dies, saying why rather than leaving
somebody to wonder what the missing thirty-seven minutes would have
said. Skipped when the caller named its own files.
It measured 164 seconds against forty-odd minutes, and failed three
times on its first run for one reason: applying the bundle raises a
control plane but does not tell it a machine exists. I had left out
enrolment, and the long suite would have taken forty minutes to say so.
The danger is not that the suite breaks. It is that nobody notices it
stopped running (novox/hq 04-ISSUES/005). The harness this replaces had
not built for two and a half months and nothing said so — and this suite
needs a hypervisor, so it inherits exactly that: it runs when somebody
remembers, and remembering is not a mechanism.
So running, recording, and rebuilding are one act:
- the host binary, control-plane image and builder are rebuilt from
source first. The last two both parse manifests; building one and not
the other left a binary eleven hours old refusing a field the mesh had
just renamed, found by a full run.
- a receipt lands in XDG state — outside git, because the question is
whether *this machine* has run it, and a receipt in git would be a
claim about everybody's machine made by whoever committed last.
- `last-run` judges it and exits non-zero when it no longer counts.
Three faults found by running the thing rather than reading it, each now
held by a test confirmed to fail without it:
- counted() passed every test while parsing nothing. The runner colours
its summary even into a pipe; the fixtures were clean text that had
been imagined rather than captured. A fixture that agrees with the
mistake proves the mistake.
- a receipt for `suite test/lastrun.test.ts` was indistinguishable from
one for the real thing — 005's own symptom, rebuilt inside its remedy.
The receipt now records what ran.
- a tree with uncommitted work reported the bare commit, claiming
coverage of code nobody can check out. Nothing else could tell: the
hash is identical either way.
Proven on real machines: 22/22, against all three repositories.