Scenario lifecycle: raise, exec, snapshot, restore, destroy

A declaration goes in and a disposable mesh comes out. Verified on a
workstation, not asserted: two machines raised and addressed in 14.6s,
snapshot 0.28s, restore-to-usable 11.6s, both families pinging with no
loss, and the workstation with no route into any of it.

The declaration layer implements the model in full — three positions a
machine can be in, keyed on forwardability; gateways carrying the address
the world sees them as; both address families; multi-homing; MTU;
inter-segment policy. It is validated hard because the failures it prevents
are silent: a private range on a public segment produces no error, the mesh
simply never forms. Public segments are refused unless they use RFC 5737 or
RFC 3849 space, and a range wider than the reserved block is refused too.
33 tests, all offline.

The runtime implements less than the model, and refuses the difference.
A scenario declaring gateways, published ports, policy, inbound deny or
place is rejected at raise with every gap named. Raising it would produce a
mesh that silently lacks what it declared, which is the fault this lab
exists to catch — 04-ISSUES/003, where a firewall key is declared in five
manifests and read by no code.

Three bugs found by review and by running it, all of one family:

The readiness check truthiness-tested incusOk's return. `exec … true`
succeeds with EMPTY output, so every machine reported unreachable while
incus exec on it worked perfectly. succeeds() now exists so the mistake is
not available, and network delete had the same bug — it counted zero
segments removed while removing them.

list() split instance from machine on the last dash, so a machine called
home-server absorbed half the instance id and destroy found nothing.
Resources are now found by the metadata they carry, never by name.

restore reported success in 0.79s while the machine's agent was still
starting, so the next command failed. Both raise and restore now wait for
usable and say how long that took — reporting the earlier number is
transport reported as effect, which is the fault the lab is being built to
find.

Two incus behaviours worth recording. Its CLI reads a YAML definition from
stdin when stdin is not a terminal, so a spawned command hangs until the
timeout kills it and arrives with empty stderr — a failure with no
explanation, on a command that works when typed. And it assigns a MAC at
runtime without recording it in device config, so MACs are derived and set
explicitly, which the guest needs anyway: it names interfaces by bus
position, and matching by name configures the wrong one on a multi-homed
machine.

No build step; Node strips the types. The lifecycle has no unit tests
because a fake hypervisor would assert that the fake behaves as expected,
which is the shape of test this project exists to stop shipping.
This commit is contained in:
2026-08-24 01:12:49 +02:00
parent 021ef4a5d7
commit a27d861d3b
24 changed files with 2331 additions and 22 deletions
+153
View File
@@ -0,0 +1,153 @@
/**
* What happens to a scenario once it is raised: inspect it, run things in it, capture and
* return it to a state, and tear it down.
*
* Snapshots are WHOLE-SCENARIO. Per-machine would be cheaper and wrong: the mesh keeps
* state that spans nodes, so restoring one machine to an earlier moment while its peers
* move on produces a mesh that has never existed and could not. Faults found there would
* be artefacts of the lab.
*/
import { incus, incusOk, succeeds, taggedInstances, taggedNetworks } from "../incus/client.ts";
import { machineName } from "./names.ts";
import { waitUntilAllUsable } from "./ready.ts";
export interface Instance {
instanceId: string;
machines: { name: string; machine: string; status: string }[];
}
/**
* Every scenario instance the daemon currently holds, found by the metadata each resource
* carries rather than by parsing names — a machine called `home-server` would otherwise
* be split in the wrong place and its instance would appear not to exist.
*/
export async function list(): Promise<Instance[]> {
const byInstance = new Map<string, Instance["machines"]>();
for (const item of await taggedInstances()) {
const entry = byInstance.get(item.instanceId) ?? [];
entry.push({ name: item.name, machine: item.machine, status: item.status });
byInstance.set(item.instanceId, entry);
}
return [...byInstance.entries()]
.map(([instanceId, machines]) => ({
instanceId,
machines: machines.sort((a, b) => a.machine.localeCompare(b.machine)),
}))
.sort((a, b) => a.instanceId.localeCompare(b.instanceId));
}
async function machinesOf(instanceId: string): Promise<string[]> {
const found = (await list()).find((i) => i.instanceId === instanceId);
if (!found) throw new Error(`no scenario instance '${instanceId}'`);
return found.machines.map((m) => m.name);
}
/**
* Run a command on a machine, through incus rather than over IP.
*
* A reachability question is therefore asked from INSIDE: *can this machine reach that
* one* is exec on the first, testing the second. The workstation is not on the scenario's
* network and its opinion would be a different question with a similar-looking answer.
*/
export async function exec(
instanceId: string,
machine: string,
command: string[],
): Promise<{ stdout: string; stderr: string }> {
const found = (await taggedInstances()).find(
(i) => i.instanceId === instanceId && i.machine === machine,
);
const name = found?.name ?? machineName(instanceId, machine);
return incus(["exec", name, "--", ...command], 120_000);
}
/**
* Capture the whole scenario as one state. Every machine, one name.
*
* Machines are snapshotted while running, so what is captured is the disk and not memory —
* crash-consistent rather than a paused mesh. Whether a mesh restored that way is coherent
* is an open question in the design, not something this silently assumes away.
*/
export async function snapshot(instanceId: string, label: string): Promise<number> {
const machines = await machinesOf(instanceId);
const started = Date.now();
for (const name of machines) {
await incus(["snapshot", "create", name, label], 300_000);
}
return (Date.now() - started) / 1000;
}
export interface RestoreResult {
/** How long the restore itself took. */
restoreSeconds: number;
/** How long until the scenario was usable again — the number that matters. */
usableSeconds: number;
}
/**
* Return the whole scenario to a state. Restoring a subset would produce a mesh that never
* was, so this is all-or-nothing.
*
* Restoring a virtual machine replaces its disk and the machine comes back up, so it
* reports RUNNING while its agent is still starting — measured, the restore call returns
* in 0.79s and the very next command fails. Reporting that as "restored" would be
* transport reported as effect, so this waits for usable and returns both numbers.
*/
export async function restore(
instanceId: string,
label: string,
readyTimeoutSeconds = 180,
log: (message: string) => void = () => {},
): Promise<RestoreResult> {
const machines = await machinesOf(instanceId);
const started = Date.now();
for (const name of machines) {
await incus(["snapshot", "restore", name, label], 300_000);
}
const restoreSeconds = (Date.now() - started) / 1000;
for (const name of machines) {
await succeeds(["start", name], 60_000);
}
await waitUntilAllUsable(machines, readyTimeoutSeconds, log);
return { restoreSeconds, usableSeconds: (Date.now() - started) / 1000 };
}
export async function snapshots(instanceId: string): Promise<string[]> {
const machines = await machinesOf(instanceId);
const first = machines[0];
if (!first) return [];
const csv = (await incusOk(["snapshot", "list", first, "--format", "csv"])) ?? "";
return csv
.split("\n")
.filter(Boolean)
.map((line) => line.split(",")[0] ?? "")
.filter(Boolean);
}
/**
* Tear the instance down: machines first, then the links they were on.
*
* Networks are removed last and only if empty — a link still carrying an interface
* cannot be deleted, and forcing it would leave the daemon with a reference to something
* gone.
*/
export async function destroy(instanceId: string): Promise<{ machines: number; networks: number }> {
const machines = (await taggedInstances()).filter((i) => i.instanceId === instanceId);
for (const machine of machines) {
await succeeds(["delete", "--force", machine.name], 300_000);
}
// Links go last and only once nothing is attached: a network still carrying an
// interface cannot be deleted, and forcing it would leave a dangling reference.
let networks = 0;
for (const network of await taggedNetworks()) {
if (network.instanceId !== instanceId) continue;
// `network delete` also succeeds silently — counted with succeeds(), not truthiness.
if (await succeeds(["network", "delete", network.name], 30_000)) networks++;
}
return { machines: machines.length, networks };
}