The lab had a registry that production does not, so it tested a fiction

The lab raised a `registry` VM, pushed ~73 images into it from the workstation, and
rewrote every manifest reference — third-party ones included — to point at it. No
production mesh has such a thing. So every bed proved that a machine could fetch an
image from a registry that exists nowhere else, and the bootstrap problems that only
appear when a machine has to fetch for itself went unfound.

What replaces it is the two things that are true in the world:

**Public images come from the public internet.** mesh-lab already created a NAT'd
uplink for exactly this and attached it to any machine declaring `egress`; no scenario
ever declared it. They do now, and third-party references are left exactly as the
catalogue writes them.

**The mesh's own images have no registry and never will.** mesh-control, mesh-builder,
mesh-route-proxy and the per-module runtimes are built from source and exist in no
registry. A machine gets them the way an operator's machine does — they are built here
and loaded onto it — and is then named by the digest of its own image configuration,
which mesh-host now accepts as "an image this machine already holds".

`images:` therefore means only *ours*, and a third-party entry is refused rather than
quietly loaded: otherwise the fiction returns one convenient line at a time. It is
per-machine as well, because "everything, everywhere" was never a description of
anything real — handing whole-mesh-full's union to its two 30GiB workstations would
fill the disk with runtimes nothing on them will start.

**The uplink and the declared gateway would have fought, silently.** A gateway container
and the transit router reach the scenario and nothing else; a default route through
either is a black hole for anything outside, and it beats the uplink's DHCP route on
metric. So a machine with egress states the scenario's ranges explicitly — through the
same gateway or transit it would have defaulted to, so the overlay-across-NAT path is
unchanged — and leaves the default to the uplink. A range with no path inside the
scenario becomes `unreachable` rather than falling through: 192.168.1.0/24 is an
ordinary private range in fact, and letting it escape would put scenario traffic on
whatever network the workstation is sitting on. `scenarioRoutesFor` is pure and tested,
because a decision only a full raise could check is one nobody checks.

The registry-reachability check the raise gained earlier is kept, pointed at the real
thing: every machine with egress must resolve a name and reach the internet before the
raise says it finished. Same failure it was written for — a raise that returns, an apply
that dies on its first pull, an instance left a bare shell — now guarding the path that
actually carries.

The base image's trust of the documentation ranges as plain-HTTP registries STAYS. It
was never only for the lab's registry: the mesh has one of its own, the `registry`
module, serving artifacts to the whole mesh over plain HTTP from whatever node runs it.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
This commit is contained in:
2026-09-10 23:16:05 +02:00
parent 0a0c57b610
commit 751948f0f9
13 changed files with 558 additions and 602 deletions
+25 -4
View File
@@ -110,6 +110,19 @@ export interface Machine {
*/
memory?: string;
cpus?: number;
/**
* Which of the scenario's `images:` this machine is handed.
*
* Absent means all of them, which is right for a one-machine bed and wrong for a mesh: a
* workstation running one small module does not want forty runtimes copied onto a 30GiB disk.
* That is not a lab economy, it is what is true — an operator's machine holds the images its own
* modules need, because somebody put them there.
*
* Every entry must appear in the scenario's `images:`. Naming one that does not is refused
* rather than ignored, because a machine silently missing an image fails much later, inside an
* apply, as a container that will not start.
*/
images?: string[];
/**
* Root disk size, e.g. "60GiB". Left unset, the VM uses the storage pool's default, which is
* enough for a handful of modules. A broad install that stocks many runtime + service images
@@ -142,11 +155,19 @@ export interface Scenario {
policy?: Policy[];
place?: Placement;
/**
* Container images this scenario needs inside it.
* **The mesh's own images** — the ones that exist in no registry and are put onto a machine by
* whoever built them.
*
* A sealed machine cannot reach a registry, so the lab raises one on a public segment and
* serves these from it. Written as tags — the digest a declaration pins is the one THIS
* registry assigns, and it is reported when the scenario is raised.
* mesh-control, mesh-builder, mesh-route-proxy, the per-module runtimes and the provisioners are
* built from source and published nowhere. A machine gets them the way an operator's machine
* does: they are built on the workstation, loaded onto the machine, and named by the digest of
* their own image configuration. Written as tags, because a tag is what `docker save` can
* export; what a declaration then pins is the image ID, reported when the scenario is raised.
*
* **Third-party images do not belong here.** postgres, gitea, the mailu stack and everything
* else are pulled from the internet over a machine's `egress` uplink, exactly as they are in
* production. The lab used to serve them from a registry of its own, and that registry did not
* exist anywhere else — so every bootstrap problem it papered over went unfound.
*/
images?: string[];
/** Name the state once placement finishes, so a run can return to it. */