The lab raised a `registry` VM, pushed ~73 images into it from the workstation, and rewrote every manifest reference — third-party ones included — to point at it. No production mesh has such a thing. So every bed proved that a machine could fetch an image from a registry that exists nowhere else, and the bootstrap problems that only appear when a machine has to fetch for itself went unfound. What replaces it is the two things that are true in the world: **Public images come from the public internet.** mesh-lab already created a NAT'd uplink for exactly this and attached it to any machine declaring `egress`; no scenario ever declared it. They do now, and third-party references are left exactly as the catalogue writes them. **The mesh's own images have no registry and never will.** mesh-control, mesh-builder, mesh-route-proxy and the per-module runtimes are built from source and exist in no registry. A machine gets them the way an operator's machine does — they are built here and loaded onto it — and is then named by the digest of its own image configuration, which mesh-host now accepts as "an image this machine already holds". `images:` therefore means only *ours*, and a third-party entry is refused rather than quietly loaded: otherwise the fiction returns one convenient line at a time. It is per-machine as well, because "everything, everywhere" was never a description of anything real — handing whole-mesh-full's union to its two 30GiB workstations would fill the disk with runtimes nothing on them will start. **The uplink and the declared gateway would have fought, silently.** A gateway container and the transit router reach the scenario and nothing else; a default route through either is a black hole for anything outside, and it beats the uplink's DHCP route on metric. So a machine with egress states the scenario's ranges explicitly — through the same gateway or transit it would have defaulted to, so the overlay-across-NAT path is unchanged — and leaves the default to the uplink. A range with no path inside the scenario becomes `unreachable` rather than falling through: 192.168.1.0/24 is an ordinary private range in fact, and letting it escape would put scenario traffic on whatever network the workstation is sitting on. `scenarioRoutesFor` is pure and tested, because a decision only a full raise could check is one nobody checks. The registry-reachability check the raise gained earlier is kept, pointed at the real thing: every machine with egress must resolve a name and reach the internet before the raise says it finished. Same failure it was written for — a raise that returns, an apply that dies on its first pull, an instance left a bare shell — now guarding the path that actually carries. The base image's trust of the documentation ranges as plain-HTTP registries STAYS. It was never only for the lab's registry: the mesh has one of its own, the `registry` module, serving artifacts to the whole mesh over plain HTTP from whatever node runs it. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
176 lines
7.5 KiB
TypeScript
176 lines
7.5 KiB
TypeScript
/**
|
|
* A scenario declares an UNDERLAY and what to place on it — the facts a machine would
|
|
* have before any of our software touched it. It declares nothing the mesh is
|
|
* responsible for: no overlay addresses, no hub, no peering, no names, no certificates.
|
|
* Those are outcomes to observe, and a scenario that supplied them would be certifying
|
|
* its own work.
|
|
*
|
|
* See novox/hq: 02-DECISIONS/0031-the-lab-provides-the-underlay.md
|
|
* 03-DESIGN/01-to-be/02-scenario-declaration.md
|
|
*/
|
|
|
|
/** An IP family. Reachability is a property of (machine, family), never of a machine. */
|
|
export type Family = "v4" | "v6";
|
|
|
|
/**
|
|
* How a segment reaches its parent.
|
|
*
|
|
* `address` is the address the outside world sees the network as — for a household
|
|
* connection, what the ISP hands out. It is load-bearing rather than decorative: it is
|
|
* what a peer records as an endpoint when a machine here dials out, and what a public
|
|
* name for a published machine here resolves to.
|
|
*/
|
|
export interface Gateway {
|
|
/** Parent segment name. */
|
|
to: string;
|
|
/** Addresses the gateway holds on the parent segment, one per family. */
|
|
address: string[];
|
|
/**
|
|
* Which families are translated. `["v4"]` is the modern default — v4 translated, v6
|
|
* routed. `[]` is a routed range where machines keep their own addresses.
|
|
*/
|
|
nat: Family[];
|
|
/**
|
|
* Whether an inbound mapping can be created. Independent of `nat`, and the field that
|
|
* separates a home gateway from carrier-grade NAT — which is your own connection and
|
|
* still unforwardable.
|
|
*/
|
|
forwardable: boolean;
|
|
/**
|
|
* How long an unused inbound mapping survives, e.g. "120s". Absent means mappings never
|
|
* expire, which no real gateway does — so absence is a simplification, not a default.
|
|
*/
|
|
mappingTtl?: string;
|
|
}
|
|
|
|
/** A broadcast domain. Several public segments are unrelated and routed, never bridged. */
|
|
export interface Segment {
|
|
/**
|
|
* `public` stands in for a public network — and there is normally more than one,
|
|
* unrelated to each other. `private` is everything else; a private segment with no
|
|
* gateway is an island that reaches nothing.
|
|
*/
|
|
kind: "public" | "private";
|
|
/** Address ranges, one per family. */
|
|
cidr: string[];
|
|
/** Largest packet the segment carries. Default 1500. Lower reproduces tunnelled paths. */
|
|
mtu?: number;
|
|
gateway?: Gateway;
|
|
}
|
|
|
|
/** Where a machine sits: a segment and the addresses it holds there. */
|
|
export interface Attachment {
|
|
segment: string;
|
|
address: string[];
|
|
}
|
|
|
|
/** A destination-NAT rule on a named gateway, stated as an outcome rather than a port list. */
|
|
export interface Publication {
|
|
port: number;
|
|
/** The segment whose gateway forwards. Named, because a machine may sit behind several. */
|
|
on: string;
|
|
}
|
|
|
|
export interface Machine {
|
|
/**
|
|
* One attachment, or several for a machine on multiple segments at once. Multi-homing
|
|
* is not exotic: it is what any node with both a LAN and a WAN interface is.
|
|
* `"detached"` is a machine on no segment — it exists and reaches nothing.
|
|
*/
|
|
at: Attachment[] | "detached";
|
|
published?: Publication[];
|
|
/**
|
|
* A host firewall. Distinct from NAT and behaves differently: a machine can be perfectly
|
|
* routable and still refuse everything unsolicited, which is the normal state of a
|
|
* v6-addressed machine. Without this, v6 addressing would imply reachability.
|
|
*/
|
|
inbound?: "allow" | "deny";
|
|
/**
|
|
* Whether this machine can reach the world outside the scenario.
|
|
*
|
|
* **Off unless asked for.** A scenario is a closed address space, and a machine that could
|
|
* reach anything would make every test's result depend on what else was reachable that day.
|
|
* It is declared for the same reason an address is: so what a run proves is what the scenario
|
|
* says, and not what the workstation happened to have.
|
|
*
|
|
* What it is for is the one thing a mesh genuinely cannot do without an outside: a first node
|
|
* fetching the images it starts from, before there is any mesh to serve them
|
|
* (novox/hq 04-ISSUES/029).
|
|
*/
|
|
egress?: boolean;
|
|
/**
|
|
* How big the machine is. Absent means the lab's default, which suits a machine running a host
|
|
* and a handful of containers.
|
|
*
|
|
* Declared, because it is a fact about the machine the scenario describes — the node that runs
|
|
* the whole substrate is bigger than the laptop that joins it, and a test that starves its
|
|
* anchor at the default answers questions about memory pressure, not about the mesh. The forge
|
|
* test failed three times as "status hangs" before anyone counted the containers in 1GiB
|
|
* (novox/hq 04-ISSUES/024 is the same lesson about a different resource).
|
|
*/
|
|
memory?: string;
|
|
cpus?: number;
|
|
/**
|
|
* Which of the scenario's `images:` this machine is handed.
|
|
*
|
|
* Absent means all of them, which is right for a one-machine bed and wrong for a mesh: a
|
|
* workstation running one small module does not want forty runtimes copied onto a 30GiB disk.
|
|
* That is not a lab economy, it is what is true — an operator's machine holds the images its own
|
|
* modules need, because somebody put them there.
|
|
*
|
|
* Every entry must appear in the scenario's `images:`. Naming one that does not is refused
|
|
* rather than ignored, because a machine silently missing an image fails much later, inside an
|
|
* apply, as a container that will not start.
|
|
*/
|
|
images?: string[];
|
|
/**
|
|
* Root disk size, e.g. "60GiB". Left unset, the VM uses the storage pool's default, which is
|
|
* enough for a handful of modules. A broad install that stocks many runtime + service images
|
|
* (each hundreds of MB, the heavy app images over a GB) exhausts the default and the host fails
|
|
* mid-apply with "no space left on device" — a disk fact about the machine, not a mesh defect.
|
|
*/
|
|
disk?: string;
|
|
}
|
|
|
|
/** Reachability between segments, as a segmented router enforces it. Asymmetric by design. */
|
|
export interface Policy {
|
|
from: string;
|
|
to: string;
|
|
allow: boolean;
|
|
}
|
|
|
|
/** What goes inside the machines. The ONLY part that differs between scenario classes. */
|
|
export interface Placement {
|
|
/** Applied to every machine. */
|
|
all?: string[];
|
|
/** Per-machine, overriding `all` for that machine. */
|
|
[machine: string]: string[] | undefined;
|
|
}
|
|
|
|
export interface Scenario {
|
|
/** The kind. Instances are many; this names the shape, not one of them. */
|
|
scenario: string;
|
|
segments: Record<string, Segment>;
|
|
machines: Record<string, Machine>;
|
|
policy?: Policy[];
|
|
place?: Placement;
|
|
/**
|
|
* **The mesh's own images** — the ones that exist in no registry and are put onto a machine by
|
|
* whoever built them.
|
|
*
|
|
* mesh-control, mesh-builder, mesh-route-proxy, the per-module runtimes and the provisioners are
|
|
* built from source and published nowhere. A machine gets them the way an operator's machine
|
|
* does: they are built on the workstation, loaded onto the machine, and named by the digest of
|
|
* their own image configuration. Written as tags, because a tag is what `docker save` can
|
|
* export; what a declaration then pins is the image ID, reported when the scenario is raised.
|
|
*
|
|
* **Third-party images do not belong here.** postgres, gitea, the mailu stack and everything
|
|
* else are pulled from the internet over a machine's `egress` uplink, exactly as they are in
|
|
* production. The lab used to serve them from a registry of its own, and that registry did not
|
|
* exist anywhere else — so every bootstrap problem it papered over went unfound.
|
|
*/
|
|
images?: string[];
|
|
/** Name the state once placement finishes, so a run can return to it. */
|
|
snapshot?: string;
|
|
}
|