The lab had a registry that production does not, so it tested a fiction
The lab raised a `registry` VM, pushed ~73 images into it from the workstation, and rewrote every manifest reference — third-party ones included — to point at it. No production mesh has such a thing. So every bed proved that a machine could fetch an image from a registry that exists nowhere else, and the bootstrap problems that only appear when a machine has to fetch for itself went unfound. What replaces it is the two things that are true in the world: **Public images come from the public internet.** mesh-lab already created a NAT'd uplink for exactly this and attached it to any machine declaring `egress`; no scenario ever declared it. They do now, and third-party references are left exactly as the catalogue writes them. **The mesh's own images have no registry and never will.** mesh-control, mesh-builder, mesh-route-proxy and the per-module runtimes are built from source and exist in no registry. A machine gets them the way an operator's machine does — they are built here and loaded onto it — and is then named by the digest of its own image configuration, which mesh-host now accepts as "an image this machine already holds". `images:` therefore means only *ours*, and a third-party entry is refused rather than quietly loaded: otherwise the fiction returns one convenient line at a time. It is per-machine as well, because "everything, everywhere" was never a description of anything real — handing whole-mesh-full's union to its two 30GiB workstations would fill the disk with runtimes nothing on them will start. **The uplink and the declared gateway would have fought, silently.** A gateway container and the transit router reach the scenario and nothing else; a default route through either is a black hole for anything outside, and it beats the uplink's DHCP route on metric. So a machine with egress states the scenario's ranges explicitly — through the same gateway or transit it would have defaulted to, so the overlay-across-NAT path is unchanged — and leaves the default to the uplink. A range with no path inside the scenario becomes `unreachable` rather than falling through: 192.168.1.0/24 is an ordinary private range in fact, and letting it escape would put scenario traffic on whatever network the workstation is sitting on. `scenarioRoutesFor` is pure and tested, because a decision only a full raise could check is one nobody checks. The registry-reachability check the raise gained earlier is kept, pointed at the real thing: every machine with egress must resolve a name and reach the internet before the raise says it finished. Same failure it was written for — a raise that returns, an apply that dies on its first pull, an instance left a bare shell — now guarding the path that actually carries. The base image's trust of the documentation ranges as plain-HTTP registries STAYS. It was never only for the lab's registry: the mesh has one of its own, the `registry` module, serving artifacts to the whole mesh over plain HTTP from whatever node runs it. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
This commit is contained in:
@@ -110,6 +110,19 @@ export interface Machine {
|
||||
*/
|
||||
memory?: string;
|
||||
cpus?: number;
|
||||
/**
|
||||
* Which of the scenario's `images:` this machine is handed.
|
||||
*
|
||||
* Absent means all of them, which is right for a one-machine bed and wrong for a mesh: a
|
||||
* workstation running one small module does not want forty runtimes copied onto a 30GiB disk.
|
||||
* That is not a lab economy, it is what is true — an operator's machine holds the images its own
|
||||
* modules need, because somebody put them there.
|
||||
*
|
||||
* Every entry must appear in the scenario's `images:`. Naming one that does not is refused
|
||||
* rather than ignored, because a machine silently missing an image fails much later, inside an
|
||||
* apply, as a container that will not start.
|
||||
*/
|
||||
images?: string[];
|
||||
/**
|
||||
* Root disk size, e.g. "60GiB". Left unset, the VM uses the storage pool's default, which is
|
||||
* enough for a handful of modules. A broad install that stocks many runtime + service images
|
||||
@@ -142,11 +155,19 @@ export interface Scenario {
|
||||
policy?: Policy[];
|
||||
place?: Placement;
|
||||
/**
|
||||
* Container images this scenario needs inside it.
|
||||
* **The mesh's own images** — the ones that exist in no registry and are put onto a machine by
|
||||
* whoever built them.
|
||||
*
|
||||
* A sealed machine cannot reach a registry, so the lab raises one on a public segment and
|
||||
* serves these from it. Written as tags — the digest a declaration pins is the one THIS
|
||||
* registry assigns, and it is reported when the scenario is raised.
|
||||
* mesh-control, mesh-builder, mesh-route-proxy, the per-module runtimes and the provisioners are
|
||||
* built from source and published nowhere. A machine gets them the way an operator's machine
|
||||
* does: they are built on the workstation, loaded onto the machine, and named by the digest of
|
||||
* their own image configuration. Written as tags, because a tag is what `docker save` can
|
||||
* export; what a declaration then pins is the image ID, reported when the scenario is raised.
|
||||
*
|
||||
* **Third-party images do not belong here.** postgres, gitea, the mailu stack and everything
|
||||
* else are pulled from the internet over a machine's `egress` uplink, exactly as they are in
|
||||
* production. The lab used to serve them from a registry of its own, and that registry did not
|
||||
* exist anywhere else — so every bootstrap problem it papered over went unfound.
|
||||
*/
|
||||
images?: string[];
|
||||
/** Name the state once placement finishes, so a run can return to it. */
|
||||
|
||||
Reference in New Issue
Block a user