Scenario lifecycle: raise, exec, snapshot, restore, destroy
A declaration goes in and a disposable mesh comes out. Verified on a workstation, not asserted: two machines raised and addressed in 14.6s, snapshot 0.28s, restore-to-usable 11.6s, both families pinging with no loss, and the workstation with no route into any of it. The declaration layer implements the model in full — three positions a machine can be in, keyed on forwardability; gateways carrying the address the world sees them as; both address families; multi-homing; MTU; inter-segment policy. It is validated hard because the failures it prevents are silent: a private range on a public segment produces no error, the mesh simply never forms. Public segments are refused unless they use RFC 5737 or RFC 3849 space, and a range wider than the reserved block is refused too. 33 tests, all offline. The runtime implements less than the model, and refuses the difference. A scenario declaring gateways, published ports, policy, inbound deny or place is rejected at raise with every gap named. Raising it would produce a mesh that silently lacks what it declared, which is the fault this lab exists to catch — 04-ISSUES/003, where a firewall key is declared in five manifests and read by no code. Three bugs found by review and by running it, all of one family: The readiness check truthiness-tested incusOk's return. `exec … true` succeeds with EMPTY output, so every machine reported unreachable while incus exec on it worked perfectly. succeeds() now exists so the mistake is not available, and network delete had the same bug — it counted zero segments removed while removing them. list() split instance from machine on the last dash, so a machine called home-server absorbed half the instance id and destroy found nothing. Resources are now found by the metadata they carry, never by name. restore reported success in 0.79s while the machine's agent was still starting, so the next command failed. Both raise and restore now wait for usable and say how long that took — reporting the earlier number is transport reported as effect, which is the fault the lab is being built to find. Two incus behaviours worth recording. Its CLI reads a YAML definition from stdin when stdin is not a terminal, so a spawned command hangs until the timeout kills it and arrives with empty stderr — a failure with no explanation, on a command that works when typed. And it assigns a MAC at runtime without recording it in device config, so MACs are derived and set explicitly, which the guest needs anyway: it names interfaces by bus position, and matching by name configures the wrong one on a multi-homed machine. No build step; Node strips the types. The lifecycle has no unit tests because a fake hypervisor would assert that the fake behaves as expected, which is the shape of test this project exists to stop shipping.
This commit is contained in:
@@ -0,0 +1,202 @@
|
||||
/**
|
||||
* Materialise a declaration into a running scenario instance.
|
||||
*
|
||||
* Two rules from the design shape everything here.
|
||||
*
|
||||
* `raise` waits for the machines to be USABLE, not for the calls to return. Measured on a
|
||||
* workstation those are 3.4s and 14.3s apart, and reporting the earlier number would be
|
||||
* the mesh's own recurring failure — transport reported as effect.
|
||||
*
|
||||
* A failed raise LEAVES THE WRECKAGE. Tearing down on failure destroys the only evidence
|
||||
* of what went wrong, and a scenario that failed to raise is more interesting than one
|
||||
* that succeeded.
|
||||
*
|
||||
* See novox/hq 03-DESIGN/01-to-be/03-scenario-lifecycle.md
|
||||
*/
|
||||
|
||||
import type { Scenario } from "../declaration/types.ts";
|
||||
import { incus, incusOk, succeeds, pools, supportedDrivers } from "../incus/client.ts";
|
||||
import { machineName, macFor, networkName, newInstanceId } from "./names.ts";
|
||||
import { waitUntilAllUsable } from "./ready.ts";
|
||||
import { applyAddresses } from "./address.ts";
|
||||
import { assertSupported } from "./supported.ts";
|
||||
|
||||
/** Drivers whose snapshots are copy-on-write. On `dir` a snapshot is a full copy. */
|
||||
const COW_DRIVERS = ["btrfs", "zfs"];
|
||||
|
||||
export interface RaiseOptions {
|
||||
/** Base image for machines. */
|
||||
image?: string;
|
||||
/** Reuse an existing instance id rather than minting one — makes raise convergent. */
|
||||
instanceId?: string;
|
||||
/** Seconds to wait for each machine to become usable. */
|
||||
readyTimeoutSeconds?: number;
|
||||
onProgress?: (message: string) => void;
|
||||
}
|
||||
|
||||
export interface RaisedScenario {
|
||||
instanceId: string;
|
||||
scenario: string;
|
||||
machines: string[];
|
||||
networks: string[];
|
||||
pool: string;
|
||||
}
|
||||
|
||||
export class RaiseError extends Error {
|
||||
readonly instanceId: string;
|
||||
readonly step: string;
|
||||
|
||||
constructor(instanceId: string, step: string, cause: unknown) {
|
||||
const detail = cause instanceof Error ? cause.message : String(cause);
|
||||
super(
|
||||
`raise failed at '${step}': ${detail}\n` +
|
||||
`The instance '${instanceId}' has been LEFT STANDING for inspection. ` +
|
||||
`Destroy it with: mesh-lab destroy ${instanceId}`,
|
||||
);
|
||||
this.name = "RaiseError";
|
||||
this.instanceId = instanceId;
|
||||
this.step = step;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Pick a pool that can snapshot cheaply, and say so loudly when there is not one.
|
||||
*
|
||||
* Measured: a `dir` snapshot of a 1.5 GB machine takes 9.9s and a full 1.6 GB, with a
|
||||
* second snapshot unfinished after two minutes. On copy-on-write it is 0.13s and costs
|
||||
* the delta. Restoring is the operation the inner loop repeats most, so a `dir` pool does
|
||||
* not make the lab slow — it makes it unused.
|
||||
*/
|
||||
async function choosePool(log: (m: string) => void): Promise<string> {
|
||||
const available = await pools();
|
||||
const cow = available.find((p) => COW_DRIVERS.includes(p.driver));
|
||||
if (cow) return cow.name;
|
||||
|
||||
const drivers = await supportedDrivers();
|
||||
const possible = drivers.filter((d) => COW_DRIVERS.includes(d));
|
||||
log(
|
||||
possible.length > 0
|
||||
? `WARNING: no copy-on-write pool exists, though the daemon offers ${possible.join("/")}. ` +
|
||||
`Snapshots will be full copies — roughly 76x slower, and the inner loop unusable.`
|
||||
: `WARNING: the daemon offers no copy-on-write driver. Snapshots will be full copies — ` +
|
||||
`roughly 76x slower, and the inner loop unusable. Install btrfs tooling and restart it.`,
|
||||
);
|
||||
const fallback = available[0];
|
||||
if (!fallback) throw new Error("no storage pool exists at all");
|
||||
return fallback.name;
|
||||
}
|
||||
|
||||
/**
|
||||
* One isolated link per segment. Nothing joins them to anything outside the instance, and
|
||||
* incus is told not to hand out addresses: a scenario declares the underlay, and letting
|
||||
* a hypervisor's DHCP assign addresses would be the lab supplying facts the declaration
|
||||
* is supposed to own.
|
||||
*/
|
||||
async function createNetwork(instanceId: string, segment: string): Promise<string> {
|
||||
const name = networkName(instanceId, segment);
|
||||
if (await succeeds(["network", "show", name], 15_000)) return name;
|
||||
await incus([
|
||||
"network", "create", name,
|
||||
"ipv4.address=none",
|
||||
"ipv6.address=none",
|
||||
"ipv4.nat=false",
|
||||
"ipv6.nat=false",
|
||||
`user.mesh-lab.instance=${instanceId}`,
|
||||
`user.mesh-lab.segment=${segment}`,
|
||||
]);
|
||||
return name;
|
||||
}
|
||||
|
||||
async function createMachine(
|
||||
instanceId: string,
|
||||
machine: string,
|
||||
attachments: { segment: string }[],
|
||||
image: string,
|
||||
pool: string,
|
||||
): Promise<string> {
|
||||
const name = machineName(instanceId, machine);
|
||||
if (await succeeds(["config", "show", name], 15_000)) return name;
|
||||
|
||||
const args = [
|
||||
"init", image, name,
|
||||
"--vm",
|
||||
"-s", pool,
|
||||
// Arch images refuse to boot under secureboot with the shipped keys. Discovered by
|
||||
// the first launch failing with exactly that message.
|
||||
"-c", "security.secureboot=false",
|
||||
"-c", "limits.memory=1GiB",
|
||||
"-c", "limits.cpu=2",
|
||||
"-c", `user.mesh-lab.instance=${instanceId}`,
|
||||
"-c", `user.mesh-lab.machine=${machine}`,
|
||||
];
|
||||
await incus(args, 300_000);
|
||||
|
||||
// eth0 comes from the profile and points at the wrong network, so every attachment is
|
||||
// explicit. A machine on no segment gets no interface at all — that is what detached is.
|
||||
await succeeds(["config", "device", "remove", name, "eth0"], 15_000);
|
||||
for (const [index, attachment] of attachments.entries()) {
|
||||
await incus([
|
||||
"config", "device", "add", name, `eth${index}`, "nic",
|
||||
"nictype=bridged",
|
||||
`parent=${networkName(instanceId, attachment.segment)}`,
|
||||
// Explicit, because incus assigns one at runtime without recording it in the device
|
||||
// config — so reading it back returns nothing, and the guest has no stable handle.
|
||||
`hwaddr=${macFor(instanceId, machine, index)}`,
|
||||
]);
|
||||
}
|
||||
return name;
|
||||
}
|
||||
|
||||
export async function raise(
|
||||
scenario: Scenario,
|
||||
options: RaiseOptions = {},
|
||||
): Promise<RaisedScenario> {
|
||||
const log = options.onProgress ?? (() => {});
|
||||
const image = options.image ?? "images:archlinux/current";
|
||||
const readyTimeout = options.readyTimeoutSeconds ?? 180;
|
||||
const instanceId = options.instanceId ?? newInstanceId(scenario.scenario, new Date());
|
||||
|
||||
// Refuse before spending a minute raising something that would silently lack half of
|
||||
// what it declares. Deliberately outside the try: this is not a raise failure, nothing
|
||||
// has been created, and there is no wreckage to leave standing.
|
||||
assertSupported(scenario);
|
||||
|
||||
let step = "choosing a storage pool";
|
||||
try {
|
||||
const pool = await choosePool(log);
|
||||
log(`instance ${instanceId} pool ${pool}`);
|
||||
|
||||
step = "creating segments";
|
||||
const networks: string[] = [];
|
||||
for (const segment of Object.keys(scenario.segments)) {
|
||||
networks.push(await createNetwork(instanceId, segment));
|
||||
log(` segment ${segment}`);
|
||||
}
|
||||
|
||||
step = "creating machines";
|
||||
const created: string[] = [];
|
||||
const byMachine = new Map<string, string>();
|
||||
for (const [machine, spec] of Object.entries(scenario.machines)) {
|
||||
const attachments = spec.at === "detached" ? [] : spec.at;
|
||||
const name = await createMachine(instanceId, machine, attachments, image, pool);
|
||||
created.push(name);
|
||||
byMachine.set(machine, name);
|
||||
log(` machine ${machine}${spec.at === "detached" ? " (detached)" : ""}`);
|
||||
}
|
||||
|
||||
step = "starting machines";
|
||||
for (const name of created) {
|
||||
await succeeds(["start", name], 60_000);
|
||||
}
|
||||
|
||||
step = "waiting for machines to become usable";
|
||||
await waitUntilAllUsable(created, readyTimeout, log);
|
||||
|
||||
step = "applying declared addresses";
|
||||
await applyAddresses(scenario, instanceId, byMachine, log);
|
||||
|
||||
return { instanceId, scenario: scenario.scenario, machines: created, networks, pool };
|
||||
} catch (cause) {
|
||||
throw new RaiseError(instanceId, step, cause);
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user