Files
mesh-lab/test/integration/harness.ts
T
jschoubben ca2bbab836 Integration tests, each named for the decision it defends
Reviewed and the criticism was right: 1,072 of 2,128 lines untested, all of
it the half that touches the hypervisor, and no gate. The verification I had
done was real — pings across NAT, TTL counts, ruleset comparisons — and none
of it survived the terminal it ran in, which is 04-ISSUES/005 in miniature.

Ten integration tests against a real hypervisor, each named for what it
defends. ADR 0031: a raised machine carries no overlay, no wireguard, no
mesh config — a scenario that pre-built peering would certify its own work.
ADR 0032: exec is the only way in. ADR 0033: routers are containers while
machines are virtual machines. And the design's claims: raise waits for
usable, snapshots are whole-scenario, NAT hides a private address,
published reaches the machine at the gateway's address.

Mocking the hypervisor is forbidden, so they skip with a reason on a
machine that cannot raise scenarios rather than passing green having
checked nothing.

The suite earned itself on its first run. It found that a snapshot of a
running machine could miss a file written seconds earlier — not stale,
absent — because the write was still in the guest's page cache. That is
exactly the question the lifecycle design listed as open: does a scenario
snapshot need the machines stopped? It does not, but it does need them
flushed. snapshot now syncs every machine before capturing, and the design
records the answer.

The fix buys write-durability, not application-consistency: anything
mid-transaction is still captured mid-transaction, and that is now stated
rather than assumed.

npm run check is the gate — typecheck, 40 unit tests, 10 integration tests.
2026-08-24 22:26:34 +02:00

45 lines
1.7 KiB
TypeScript

/**
* Integration tests run against a real hypervisor. Mocking it is forbidden — a test that
* fakes the system under integration asserts that the fake behaves as expected, which is
* the shape of test this project exists to stop shipping (novox/hq ADR 0034).
*
* Consequence, accepted: these are slow, and they need a machine that can raise scenarios.
* They skip rather than fail where it cannot, so that a machine without a hypervisor gets
* an honest "not run" instead of a green suite that checked nothing.
*/
import { isReachable, pools, supportedDrivers } from "../../src/incus/client.ts";
import { destroy, list } from "../../src/lifecycle/operate.ts";
export interface Capability {
usable: boolean;
why: string;
}
/** Can this machine run scenarios at all? Checked once, reported honestly. */
export async function labIsUsable(): Promise<Capability> {
if (!(await isReachable())) {
return {
usable: false,
why: "the incus daemon is not reachable as this user (try MESH_LAB_INCUS='sudo -n incus')",
};
}
const drivers = await supportedDrivers();
if (!drivers.some((d) => d === "btrfs" || d === "zfs")) {
return { usable: false, why: "no copy-on-write driver — snapshots would be full copies" };
}
if (!(await pools()).some((p) => p.driver === "btrfs" || p.driver === "zfs")) {
return { usable: false, why: "no pool uses a copy-on-write driver" };
}
return { usable: true, why: "" };
}
/** Tear down anything a test left behind, whether it passed or not. */
export async function destroyAll(prefix: string): Promise<void> {
for (const instance of await list()) {
if (instance.instanceId.startsWith(prefix)) {
await destroy(instance.instanceId);
}
}
}