The lab had a registry that production does not, so it tested a fiction

The lab raised a `registry` VM, pushed ~73 images into it from the workstation, and
rewrote every manifest reference — third-party ones included — to point at it. No
production mesh has such a thing. So every bed proved that a machine could fetch an
image from a registry that exists nowhere else, and the bootstrap problems that only
appear when a machine has to fetch for itself went unfound.

What replaces it is the two things that are true in the world:

**Public images come from the public internet.** mesh-lab already created a NAT'd
uplink for exactly this and attached it to any machine declaring `egress`; no scenario
ever declared it. They do now, and third-party references are left exactly as the
catalogue writes them.

**The mesh's own images have no registry and never will.** mesh-control, mesh-builder,
mesh-route-proxy and the per-module runtimes are built from source and exist in no
registry. A machine gets them the way an operator's machine does — they are built here
and loaded onto it — and is then named by the digest of its own image configuration,
which mesh-host now accepts as "an image this machine already holds".

`images:` therefore means only *ours*, and a third-party entry is refused rather than
quietly loaded: otherwise the fiction returns one convenient line at a time. It is
per-machine as well, because "everything, everywhere" was never a description of
anything real — handing whole-mesh-full's union to its two 30GiB workstations would
fill the disk with runtimes nothing on them will start.

**The uplink and the declared gateway would have fought, silently.** A gateway container
and the transit router reach the scenario and nothing else; a default route through
either is a black hole for anything outside, and it beats the uplink's DHCP route on
metric. So a machine with egress states the scenario's ranges explicitly — through the
same gateway or transit it would have defaulted to, so the overlay-across-NAT path is
unchanged — and leaves the default to the uplink. A range with no path inside the
scenario becomes `unreachable` rather than falling through: 192.168.1.0/24 is an
ordinary private range in fact, and letting it escape would put scenario traffic on
whatever network the workstation is sitting on. `scenarioRoutesFor` is pure and tested,
because a decision only a full raise could check is one nobody checks.

The registry-reachability check the raise gained earlier is kept, pointed at the real
thing: every machine with egress must resolve a name and reach the internet before the
raise says it finished. Same failure it was written for — a raise that returns, an apply
that dies on its first pull, an instance left a bare shell — now guarding the path that
actually carries.

The base image's trust of the documentation ranges as plain-HTTP registries STAYS. It
was never only for the lab's registry: the mesh has one of its own, the `registry`
module, serving artifacts to the whole mesh over plain HTTP from whatever node runs it.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
This commit is contained in:
2026-09-10 23:16:05 +02:00
parent 0a0c57b610
commit 751948f0f9
13 changed files with 558 additions and 602 deletions
+11 -10
View File
@@ -23,6 +23,7 @@ import { homedir } from "node:os";
import { list, restore, snapshot, snapshots, destroy } from "./lifecycle/operate.ts";
import type { Against } from "./lastrun.ts";
import { whatWasTested } from "./lastrun.ts";
import type { HeldImage } from "./pinning.ts";
/** The state a warm instance is kept at. One label, because a second is a state nobody named. */
export const label = "warm";
@@ -31,13 +32,13 @@ export interface Warm {
scenario: string;
instanceId: string;
/**
* The image references the scenario's registry serves, pinned by digest.
* The mesh's own images, as loaded onto this instance's machines, by the ID each is held under.
*
* Kept because they are worked out while raising and a restored instance never raises. Without
* them a warm run knows nothing about what it can pull, and every test naming an image fails
* for a reason that has nothing to do with what it was testing.
* them a warm run knows nothing about what its machines hold, and every test naming one of our
* images fails for a reason that has nothing to do with what it was testing.
*/
images: string[];
images: HeldImage[];
/** The commit each repository was at when this was brought to its state. */
against: Against;
at: string;
@@ -187,23 +188,23 @@ export async function cool(): Promise<string | null> {
}
/**
* What a raised scenario stocked, held until it is kept.
* What a raised scenario loaded onto its machines, held until it is kept.
*
* Raising works the images out and snapshotting happens later, so this carries them between the
* two without the caller having to hold them.
*/
const stock = new Map<string, string[]>();
const stock = new Map<string, HeldImage[]>();
export function rememberStock(instanceId: string, images: string[]): void {
export function rememberStock(instanceId: string, images: HeldImage[]): void {
stock.set(instanceId, images);
}
function stockOf(instanceId: string): string[] {
function stockOf(instanceId: string): HeldImage[] {
return stock.get(instanceId) ?? [];
}
/** What a restored instance's registry serves, from when it was warmed. */
export function warmStock(instanceId: string): { images: string[] } {
/** What a restored instance's machines hold, from when it was warmed. */
export function warmStock(instanceId: string): { images: HeldImage[] } {
const warm = remembered();
if (!warm || warm.instanceId !== instanceId) return { images: [] };
return { images: warm.images };