The lab raised a `registry` VM, pushed ~73 images into it from the workstation, and rewrote every manifest reference — third-party ones included — to point at it. No production mesh has such a thing. So every bed proved that a machine could fetch an image from a registry that exists nowhere else, and the bootstrap problems that only appear when a machine has to fetch for itself went unfound. What replaces it is the two things that are true in the world: **Public images come from the public internet.** mesh-lab already created a NAT'd uplink for exactly this and attached it to any machine declaring `egress`; no scenario ever declared it. They do now, and third-party references are left exactly as the catalogue writes them. **The mesh's own images have no registry and never will.** mesh-control, mesh-builder, mesh-route-proxy and the per-module runtimes are built from source and exist in no registry. A machine gets them the way an operator's machine does — they are built here and loaded onto it — and is then named by the digest of its own image configuration, which mesh-host now accepts as "an image this machine already holds". `images:` therefore means only *ours*, and a third-party entry is refused rather than quietly loaded: otherwise the fiction returns one convenient line at a time. It is per-machine as well, because "everything, everywhere" was never a description of anything real — handing whole-mesh-full's union to its two 30GiB workstations would fill the disk with runtimes nothing on them will start. **The uplink and the declared gateway would have fought, silently.** A gateway container and the transit router reach the scenario and nothing else; a default route through either is a black hole for anything outside, and it beats the uplink's DHCP route on metric. So a machine with egress states the scenario's ranges explicitly — through the same gateway or transit it would have defaulted to, so the overlay-across-NAT path is unchanged — and leaves the default to the uplink. A range with no path inside the scenario becomes `unreachable` rather than falling through: 192.168.1.0/24 is an ordinary private range in fact, and letting it escape would put scenario traffic on whatever network the workstation is sitting on. `scenarioRoutesFor` is pure and tested, because a decision only a full raise could check is one nobody checks. The registry-reachability check the raise gained earlier is kept, pointed at the real thing: every machine with egress must resolve a name and reach the internet before the raise says it finished. Same failure it was written for — a raise that returns, an apply that dies on its first pull, an instance left a bare shell — now guarding the path that actually carries. The base image's trust of the documentation ranges as plain-HTTP registries STAYS. It was never only for the lab's registry: the mesh has one of its own, the `registry` module, serving artifacts to the whole mesh over plain HTTP from whatever node runs it. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
127 lines
5.5 KiB
TypeScript
127 lines
5.5 KiB
TypeScript
/**
|
|
* Confirming a machine that says it can reach the outside actually can.
|
|
*
|
|
* **This is what the registry-reachability check became.** The old one proved that every machine
|
|
* could fetch a manifest from the registry the lab raised inside the scenario — a real check of a
|
|
* fake path, since no production mesh has such a registry. What a machine actually does is pull
|
|
* from the internet, and that is now the thing worth proving before a raise says it is finished.
|
|
*
|
|
* The failure it exists to stop is the same one, in the same shape: `raise` returns, the caller
|
|
* applies a substrate, the first pull fails, no node enrols, and the instance is left a bare
|
|
* shell — with the cause several steps back and looking like a mesh fault rather than a lab one.
|
|
*
|
|
* Two things are checked, in this order, because they fail differently and the difference is the
|
|
* whole diagnosis:
|
|
*
|
|
* - **A name resolves.** Without this the machine has a route and no way to use it, and every
|
|
* pull dies inside the runtime saying it cannot look up a host.
|
|
* - **The path carries.** A request to the registry every image ultimately comes from, over the
|
|
* uplink, through whatever gateway sits in front of this machine. Any HTTP answer counts: what
|
|
* is in question is the path, not whether Docker Hub likes us.
|
|
*/
|
|
|
|
import type { Scenario } from "../declaration/types.ts";
|
|
import { incus, incusOk } from "../incus/client.ts";
|
|
|
|
export class EgressError extends Error {
|
|
constructor(message: string) {
|
|
super(message);
|
|
this.name = "EgressError";
|
|
}
|
|
}
|
|
|
|
/** The host every image is fetched through, in the end. Asked for, never pulled from, here. */
|
|
const UPSTREAM = "registry-1.docker.io";
|
|
|
|
/**
|
|
* Confirm every machine declaring `egress` can resolve and reach the outside.
|
|
*
|
|
* Run after the routes and the firewalls, because that is the path a pull will take: a home node's
|
|
* default is the uplink, its route to the rest of the scenario is through its gateway, and its own
|
|
* filtering is in place. Checking earlier would prove something no pull relies on.
|
|
*/
|
|
export async function confirmEgress(
|
|
scenario: Scenario,
|
|
machineNames: Map<string, string>,
|
|
log: (message: string) => void = () => {},
|
|
waitSeconds = 120,
|
|
): Promise<void> {
|
|
for (const [machine, spec] of Object.entries(scenario.machines)) {
|
|
if (!spec.egress || spec.at === "detached") continue;
|
|
const name = machineNames.get(machine);
|
|
if (!name) continue;
|
|
|
|
if (!(await resolves(name, waitSeconds))) {
|
|
// One repair, then a verdict. The uplink is the lab's own network and its DHCP server is
|
|
// also its resolver, so the machine has been told the answer and may simply have nowhere
|
|
// to write it — an image without systemd-resolved leaves `UseDNS=yes` inert.
|
|
await pointResolverAtTheUplink(name);
|
|
if (!(await resolves(name, 30))) {
|
|
throw new EgressError(
|
|
`${machine} declares egress and cannot resolve ${UPSTREAM}.\n` +
|
|
` It has a route out and no way to use it, so every image pulled from the internet ` +
|
|
`would fail inside the runtime as a lookup error.\n` +
|
|
` The uplink's DHCP server is also its resolver; this machine has not taken it.`,
|
|
);
|
|
}
|
|
}
|
|
|
|
const code = await reaches(name, waitSeconds);
|
|
if (!code) {
|
|
throw new EgressError(
|
|
`${machine} declares egress, resolves names, and cannot reach ${UPSTREAM}.\n` +
|
|
` This is the PATH: its default route, the uplink, or the host's own forwarding. ` +
|
|
`Every third-party image this machine needs is pulled from the internet, so anything ` +
|
|
`applied to it would stop at the first container.`,
|
|
);
|
|
}
|
|
log(` ${machine} reaches the internet over its uplink (${UPSTREAM} answered ${code})`);
|
|
}
|
|
}
|
|
|
|
async function resolves(name: string, waitSeconds: number): Promise<boolean> {
|
|
const deadline = Date.now() + waitSeconds * 1_000;
|
|
while (Date.now() < deadline) {
|
|
const said = await incusOk(
|
|
["exec", name, "--", "sh", "-c", `getent hosts ${UPSTREAM} >/dev/null && echo yes`], 30_000,
|
|
);
|
|
if (said?.trim() === "yes") return true;
|
|
await new Promise((r) => setTimeout(r, 3_000));
|
|
}
|
|
return false;
|
|
}
|
|
|
|
/**
|
|
* Any HTTP status at all, which is what "the path carries" means.
|
|
*
|
|
* Not 200: an unauthenticated `/v2/` is answered 401 by design, and a check demanding 200 would
|
|
* fail on a machine whose network is perfect.
|
|
*/
|
|
async function reaches(name: string, waitSeconds: number): Promise<string | null> {
|
|
const deadline = Date.now() + waitSeconds * 1_000;
|
|
while (Date.now() < deadline) {
|
|
const said = (await incusOk(
|
|
["exec", name, "--", "sh", "-c",
|
|
`curl -s -o /dev/null -w '%{http_code}' --max-time 15 https://${UPSTREAM}/v2/`], 40_000,
|
|
))?.trim();
|
|
if (said && /^[1-5][0-9]{2}$/.test(said)) return said;
|
|
await new Promise((r) => setTimeout(r, 5_000));
|
|
}
|
|
return null;
|
|
}
|
|
|
|
/**
|
|
* Write a resolver of last resort: the uplink's own gateway, which serves DHCP and DNS both.
|
|
*
|
|
* Deliberately the machine's default next hop rather than a name looked up somewhere — for a
|
|
* machine with egress that is the uplink by construction, since every scenario range is routed
|
|
* explicitly and nothing else defaults.
|
|
*/
|
|
async function pointResolverAtTheUplink(name: string): Promise<void> {
|
|
await incus([
|
|
"exec", name, "--", "sh", "-c",
|
|
`via=$(ip -4 route show default | awk '{print $3}' | head -n1); ` +
|
|
`[ -n "$via" ] && printf 'nameserver %s\\n' "$via" > /etc/resolv.conf; true`,
|
|
], 30_000);
|
|
}
|