A run rebuilds what it tests, and leaves a receipt saying what it covered
The danger is not that the suite breaks. It is that nobody notices it stopped running (novox/hq 04-ISSUES/005). The harness this replaces had not built for two and a half months and nothing said so — and this suite needs a hypervisor, so it inherits exactly that: it runs when somebody remembers, and remembering is not a mechanism. So running, recording, and rebuilding are one act: - the host binary, control-plane image and builder are rebuilt from source first. The last two both parse manifests; building one and not the other left a binary eleven hours old refusing a field the mesh had just renamed, found by a full run. - a receipt lands in XDG state — outside git, because the question is whether *this machine* has run it, and a receipt in git would be a claim about everybody's machine made by whoever committed last. - `last-run` judges it and exits non-zero when it no longer counts. Three faults found by running the thing rather than reading it, each now held by a test confirmed to fail without it: - counted() passed every test while parsing nothing. The runner colours its summary even into a pipe; the fixtures were clean text that had been imagined rather than captured. A fixture that agrees with the mistake proves the mistake. - a receipt for `suite test/lastrun.test.ts` was indistinguishable from one for the real thing — 005's own symptom, rebuilt inside its remedy. The receipt now records what ran. - a tree with uncommitted work reported the bare commit, claiming coverage of code nobody can check out. Nothing else could tell: the hash is identical either way. Proven on real machines: 22/22, against all three repositories.
This commit is contained in:
+27
@@ -30,6 +30,9 @@ const USAGE = `mesh-lab — raise a disposable mesh on one machine
|
||||
diagram <scenario.yml> [out.drawio] draw what a scenario asks for
|
||||
diagram --live <instance> [out.drawio] draw what is actually raised
|
||||
|
||||
suite [paths...] [--no-build] rebuild the artifacts, run the end-to-end tests, leave a receipt
|
||||
last-run whether the last run still counts; non-zero when it does not
|
||||
|
||||
Set MESH_LAB_INCUS if the daemon needs a different invocation, e.g. "sudo -n incus".
|
||||
`;
|
||||
|
||||
@@ -91,6 +94,30 @@ async function main(): Promise<void> {
|
||||
case "check":
|
||||
return check();
|
||||
|
||||
// Rebuilding, running, and recording that it ran are one act. Separate commands would mean a
|
||||
// run against a stale artifact, or a run nobody recorded — and both are the state
|
||||
// novox/hq 04-ISSUES/005 is about.
|
||||
case "suite": {
|
||||
const { runSuite } = await import("./suite.ts");
|
||||
process.exitCode = await runSuite(rest);
|
||||
return;
|
||||
}
|
||||
|
||||
// **What 04-ISSUES/005 says nobody was ever told.** The suite needs a machine with a
|
||||
// hypervisor, so it cannot run on every push — which means it runs when somebody remembers,
|
||||
// and remembering is not a mechanism. This asks whether the last run still means anything,
|
||||
// and exits non-zero when it does not, so a timer or a person can act on it.
|
||||
case "last-run": {
|
||||
const { read, judge, whatWasTested } = await import("./lastrun.ts");
|
||||
const said = judge(read(), new Date(), whatWasTested());
|
||||
for (const line of said.lines) console.log(line);
|
||||
if (!said.current) {
|
||||
console.log("\n `npm run check` runs it.");
|
||||
process.exitCode = 1;
|
||||
}
|
||||
return;
|
||||
}
|
||||
|
||||
case "base": {
|
||||
// `base build` exists because a sealed scenario cannot install a container runtime, and
|
||||
// the runtime has to come from somewhere with a network (novox/hq ADR 0006).
|
||||
|
||||
+208
@@ -0,0 +1,208 @@
|
||||
/**
|
||||
* When the end-to-end suite last ran, and against what.
|
||||
*
|
||||
* **The fault this exists for is not that the suite breaks — it is that nobody notices it stopped
|
||||
* running** (novox/hq 04-ISSUES/005). The harness it replaces had not built for two and a half
|
||||
* months, and nothing said so; the coverage was assumed rather than checked, and several of the
|
||||
* pipeline's most expensive defects landed inside that window.
|
||||
*
|
||||
* This suite is in a better position and the same danger: it needs a machine with a hypervisor, so
|
||||
* it cannot run on every push, which means it runs when somebody remembers. Remembering is not a
|
||||
* mechanism.
|
||||
*
|
||||
* So a run leaves a receipt, and something can be asked whether the receipt still means anything.
|
||||
* A receipt that is old, or taken against code the repositories have since moved past, is the
|
||||
* thing 005 says nobody was ever told.
|
||||
*
|
||||
* **Kept outside the repository**, because the question is *has this machine run it* rather than
|
||||
* *what is committed* — and a receipt in git would be a claim about everyone's machine made by
|
||||
* whoever committed last.
|
||||
*/
|
||||
|
||||
import { execFileSync } from "node:child_process";
|
||||
import { mkdirSync, readFileSync, writeFileSync } from "node:fs";
|
||||
import { homedir } from "node:os";
|
||||
import { dirname, join } from "node:path";
|
||||
|
||||
import { repositories } from "./repos.ts";
|
||||
|
||||
/** What a run was taken against, per repository. */
|
||||
export type Against = Record<string, string>;
|
||||
|
||||
export interface Receipt {
|
||||
/** When it finished, ISO 8601. */
|
||||
at: string;
|
||||
passed: number;
|
||||
failed: number;
|
||||
/** The commit each repository was at. Absent for anything that was not a git checkout. */
|
||||
against: Against;
|
||||
/** The test files this run was pointed at. See {@link endToEnd}. */
|
||||
ran: string[];
|
||||
}
|
||||
|
||||
/**
|
||||
* endToEnd is the file that raises real machines. A run that did not include it proved nothing
|
||||
* about the pipeline, however green it was.
|
||||
*
|
||||
* The suite takes paths, so it can be pointed at one quick file — and the receipt from that would
|
||||
* otherwise be indistinguishable from a receipt for the real thing. That is 04-ISSUES/005 again:
|
||||
* not a suite that fails, a record that says more than the run behind it.
|
||||
*/
|
||||
export const endToEnd = "test/integration/mesh.test.ts";
|
||||
|
||||
/** Where the receipt lives: XDG state, which is for exactly this — data a tool keeps between runs. */
|
||||
export function receiptPath(): string {
|
||||
const state = process.env["XDG_STATE_HOME"] ?? join(homedir(), ".local", "state");
|
||||
return join(state, "mesh-lab", "last-run.json");
|
||||
}
|
||||
|
||||
/**
|
||||
* headOf is the commit a directory's repository is at, or "" if it is not one.
|
||||
*
|
||||
* **A tree with uncommitted changes is marked, and never equal to the clean commit it sits on.**
|
||||
* The run tested what was on disk, and that is not what the commit contains — so a receipt naming
|
||||
* the bare hash would claim coverage of code nobody can check out. Nothing else could tell: the
|
||||
* hash is identical either way. That is 04-ISSUES/005's overclaim in its quietest form.
|
||||
*/
|
||||
export function headOf(directory: string): string {
|
||||
const git = (args: string[]) =>
|
||||
execFileSync("git", ["-C", directory, ...args], {
|
||||
encoding: "utf8",
|
||||
stdio: ["ignore", "pipe", "ignore"],
|
||||
});
|
||||
try {
|
||||
const head = git(["rev-parse", "--short", "HEAD"]).trim();
|
||||
const dirty = git(["status", "--porcelain"]).trim() !== "";
|
||||
return dirty ? `${head}+uncommitted` : head;
|
||||
} catch {
|
||||
// Not a checkout, or no git. Absent rather than guessed: a receipt claiming a commit it did
|
||||
// not read is worse than one that says it could not tell.
|
||||
return "";
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* whatWasTested is the repositories this run exercised, by the paths it was given.
|
||||
*
|
||||
* From the environment rather than a fixed list, because the paths are how the suite is told what
|
||||
* to run — so anything it was pointed at is something the receipt should account for, and anything
|
||||
* it was not pointed at was not tested.
|
||||
*/
|
||||
export function whatWasTested(env: NodeJS.ProcessEnv = process.env): Against {
|
||||
const against: Against = {};
|
||||
for (const [name, directory] of Object.entries(repositories(env))) {
|
||||
const head = headOf(directory);
|
||||
if (head) against[name] = head;
|
||||
}
|
||||
return against;
|
||||
}
|
||||
|
||||
/** record writes the receipt. Failures are recorded too: a run that failed still ran. */
|
||||
export function record(
|
||||
passed: number,
|
||||
failed: number,
|
||||
ran: string[],
|
||||
env = process.env,
|
||||
): Receipt {
|
||||
const receipt: Receipt = {
|
||||
at: new Date().toISOString(),
|
||||
passed,
|
||||
failed,
|
||||
against: whatWasTested(env),
|
||||
ran,
|
||||
};
|
||||
const path = receiptPath();
|
||||
mkdirSync(dirname(path), { recursive: true });
|
||||
writeFileSync(path, JSON.stringify(receipt, null, 2) + "\n");
|
||||
return receipt;
|
||||
}
|
||||
|
||||
/** read returns the receipt, or null when this machine has never run the suite. */
|
||||
export function read(): Receipt | null {
|
||||
try {
|
||||
return JSON.parse(readFileSync(receiptPath(), "utf8")) as Receipt;
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
export interface Verdict {
|
||||
/** True when the receipt still says something about the code as it stands. */
|
||||
current: boolean;
|
||||
lines: string[];
|
||||
}
|
||||
|
||||
/**
|
||||
* judge says whether the last run still means anything.
|
||||
*
|
||||
* Two ways it can stop meaning something, and they read differently: it was long ago, or the code
|
||||
* has moved since. The second is the one that matters — a suite that passed against code nobody
|
||||
* runs any more is coverage in name.
|
||||
*/
|
||||
export function judge(
|
||||
receipt: Receipt | null,
|
||||
now: Date,
|
||||
against: Against,
|
||||
staleAfterDays = 7,
|
||||
): Verdict {
|
||||
if (!receipt) {
|
||||
return {
|
||||
current: false,
|
||||
lines: [
|
||||
"this machine has never run the end-to-end suite.",
|
||||
" Nothing here has been checked end to end, which is not the same as nothing being wrong.",
|
||||
],
|
||||
};
|
||||
}
|
||||
|
||||
// A receipt written before this field existed says nothing about what it ran, and the honest
|
||||
// reading of "nothing said" is not "everything".
|
||||
const endToEndRan = (receipt.ran ?? []).some((path) => path.endsWith(endToEnd));
|
||||
|
||||
const days = (now.getTime() - Date.parse(receipt.at)) / 86_400_000;
|
||||
const lines: string[] = [];
|
||||
const outcome = receipt.failed > 0
|
||||
? `last ran ${ago(days)} and ${receipt.failed} test(s) failed`
|
||||
: `last passed ${ago(days)}, ${receipt.passed} test(s)`;
|
||||
lines.push(`the end-to-end suite ${outcome}`);
|
||||
|
||||
let moved = false;
|
||||
for (const name of Object.keys(against).sort()) {
|
||||
const then = receipt.against[name];
|
||||
const now = against[name];
|
||||
if (!then) {
|
||||
lines.push(` ${name.padEnd(14)} was not accounted for in that run`);
|
||||
moved = true;
|
||||
continue;
|
||||
}
|
||||
if (then === now) {
|
||||
lines.push(` ${name.padEnd(14)} at ${then} (unchanged)`);
|
||||
continue;
|
||||
}
|
||||
lines.push(` ${name.padEnd(14)} at ${then}, now at ${now}`);
|
||||
moved = true;
|
||||
}
|
||||
|
||||
if (!endToEndRan) {
|
||||
lines.push(` That run did not include ${endToEnd}, so it raised no machines.`);
|
||||
}
|
||||
|
||||
if (receipt.failed > 0) {
|
||||
lines.push(" Nothing has been proven end to end since.");
|
||||
} else if (moved) {
|
||||
lines.push(" What it proved was proven about code that has since changed.");
|
||||
} else if (days > staleAfterDays) {
|
||||
lines.push(` Nothing has changed since, but that was more than ${staleAfterDays} days ago.`);
|
||||
}
|
||||
|
||||
return {
|
||||
current: endToEndRan && receipt.failed === 0 && !moved && days <= staleAfterDays,
|
||||
lines,
|
||||
};
|
||||
}
|
||||
|
||||
function ago(days: number): string {
|
||||
if (days < 1 / 24) return "less than an hour ago";
|
||||
if (days < 1) return `${Math.round(days * 24)} hour(s) ago`;
|
||||
return `${Math.round(days)} day(s) ago`;
|
||||
}
|
||||
@@ -0,0 +1,85 @@
|
||||
/**
|
||||
* Rebuild what the lab runs, from source, before it runs.
|
||||
*
|
||||
* **A stale artifact reporting success against old rules is the fault this project keeps writing
|
||||
* down** (novox/hq 04-ISSUES/005). The lab consumes three artifacts from two repositories, and they
|
||||
* were rebuilt by hand, one at a time, from memory. A rename in the control plane's catalogue needs
|
||||
* both the control-plane image *and* the builder binary, because both parse manifests; rebuilding
|
||||
* one left a binary eleven hours old refusing a field the mesh had just renamed, and cost a full
|
||||
* run to find out.
|
||||
*
|
||||
* In the repository rather than in a shell script beside it, for the reason 005 is about: a step
|
||||
* that lives in somebody's terminal history runs when they remember, and remembering is not a
|
||||
* mechanism.
|
||||
*/
|
||||
|
||||
import { spawnSync } from "node:child_process";
|
||||
import { repositories } from "./repos.ts";
|
||||
|
||||
export interface Build {
|
||||
/** What it produces, for the log. */
|
||||
what: string;
|
||||
/** The repository root to run in. */
|
||||
in: string;
|
||||
argv: string[];
|
||||
env?: NodeJS.ProcessEnv;
|
||||
}
|
||||
|
||||
/**
|
||||
* planned is what must be built, given where this run has been pointed.
|
||||
*
|
||||
* Derived from the same environment the suite is configured by, so there is one place that says
|
||||
* where a repository is. A repository this run was not pointed at is not built — and, per
|
||||
* {@link whatWasTested}, is not claimed in the receipt either.
|
||||
*/
|
||||
export function planned(env: NodeJS.ProcessEnv = process.env): Build[] {
|
||||
const builds: Build[] = [];
|
||||
const where = repositories(env);
|
||||
const host = env["MESH_LAB_HOST_BINARY"];
|
||||
if (host && where["mesh-host"]) {
|
||||
builds.push({
|
||||
what: "host",
|
||||
in: where["mesh-host"],
|
||||
argv: ["go", "build", "-ldflags=-s -w -X main.builtFor=arch", "-o", host, "./cmd/mesh-host"],
|
||||
env: { CGO_ENABLED: "0" },
|
||||
});
|
||||
}
|
||||
const control = where["mesh-control"];
|
||||
if (control) {
|
||||
builds.push({ what: "control plane image", in: control, argv: ["make", "image"] });
|
||||
const builder = env["MESH_LAB_BUILDER"];
|
||||
if (builder) {
|
||||
// Both of these parse manifests. Building one and not the other is the eleven-hour-old
|
||||
// binary above, so they are one step and not two.
|
||||
builds.push({
|
||||
what: "builder",
|
||||
in: control,
|
||||
argv: ["go", "build", "-o", builder, "./cmd/mesh-builder"],
|
||||
});
|
||||
}
|
||||
}
|
||||
return builds;
|
||||
}
|
||||
|
||||
/** rebuild runs the plan, and throws on the first failure rather than testing a stale artifact. */
|
||||
export function rebuild(env: NodeJS.ProcessEnv = process.env): string[] {
|
||||
const built: string[] = [];
|
||||
for (const build of planned(env)) {
|
||||
const [command, ...args] = build.argv;
|
||||
const ran = spawnSync(command!, args, {
|
||||
cwd: build.in,
|
||||
env: { ...env, ...build.env },
|
||||
encoding: "utf8",
|
||||
});
|
||||
if (ran.status !== 0) {
|
||||
// Loudly, and stopping. A suite that runs anyway is a suite reporting on code that is not
|
||||
// the code in front of you, which is the whole of 005.
|
||||
throw new Error(
|
||||
`could not build the ${build.what}: ${build.argv.join(" ")} in ${build.in}\n\n` +
|
||||
`${(ran.stderr || ran.stdout || String(ran.error)).trim()}`,
|
||||
);
|
||||
}
|
||||
built.push(build.what);
|
||||
}
|
||||
return built;
|
||||
}
|
||||
@@ -0,0 +1,28 @@
|
||||
/**
|
||||
* Where the repositories this run was pointed at are.
|
||||
*
|
||||
* **One place**, because two things need it and they must agree: the rebuild builds these, and the
|
||||
* receipt claims these. A receipt naming a repository the run did not build is exactly the
|
||||
* false coverage novox/hq 04-ISSUES/005 is about, and it would arrive by nobody's decision — just
|
||||
* two derivations drifting.
|
||||
*
|
||||
* Derived from the environment the suite is configured by, so a repository this run was not
|
||||
* pointed at is neither built nor claimed.
|
||||
*/
|
||||
|
||||
import { dirname } from "node:path";
|
||||
|
||||
export interface Repositories {
|
||||
/** Absolute path to the repository root, by name. */
|
||||
[name: string]: string;
|
||||
}
|
||||
|
||||
export function repositories(env: NodeJS.ProcessEnv = process.env): Repositories {
|
||||
const found: Repositories = {};
|
||||
found["mesh-lab"] = process.cwd();
|
||||
const host = env["MESH_LAB_HOST_BINARY"];
|
||||
if (host) found["mesh-host"] = dirname(host);
|
||||
const modules = env["MESH_LAB_MODULES"];
|
||||
if (modules) found["mesh-control"] = dirname(dirname(modules));
|
||||
return found;
|
||||
}
|
||||
@@ -0,0 +1,89 @@
|
||||
/**
|
||||
* Run the end-to-end suite, and leave a receipt saying it ran.
|
||||
*
|
||||
* **Here rather than inside the tests, because the tests cannot know their own totals.** Node's
|
||||
* runner reports them to whatever invoked it, and a test file inventing its own count would be a
|
||||
* receipt that says whatever the last edit made it say.
|
||||
*
|
||||
* Here rather than in a shell script for the same reason the rebuild is: a step that lives in
|
||||
* somebody's terminal history is a step that runs when they remember (novox/hq 04-ISSUES/005).
|
||||
*/
|
||||
|
||||
import { spawn } from "node:child_process";
|
||||
import { endToEnd, record } from "./lastrun.ts";
|
||||
import { rebuild } from "./rebuild.ts";
|
||||
|
||||
/** counted is what the runner said, or nulls when it said nothing recognisable. */
|
||||
export function counted(output: string): { passed: number | null; failed: number | null } {
|
||||
// The runner's own summary lines, each on a line of its own. Anchored, so a test *named*
|
||||
// "pass 3" cannot be mistaken for the total — which is not a hypothetical worry in a suite whose
|
||||
// tests are named in sentences.
|
||||
// Stripped first: the runner colours its summary even when its stdout is a pipe, so the line is
|
||||
// "\x1b[34m\u2139 pass 8\x1b[39m" and an anchored pattern never sees the start of it. Found by
|
||||
// running this against the real runner — the fixture it was first written against was output I
|
||||
// had imagined, which is a test that agrees with the mistake it was written beside.
|
||||
const plain = output.replace(/\u001b\[[0-9;]*m/g, "");
|
||||
const total = (what: RegExp) => {
|
||||
const found = plain.match(what);
|
||||
return found ? Number(found[1]) : null;
|
||||
};
|
||||
return {
|
||||
passed: total(/^\s*(?:\u2139|#)\s*pass\s+(\d+)\s*$/m),
|
||||
failed: total(/^\s*(?:\u2139|#)\s*fail\s+(\d+)\s*$/m),
|
||||
};
|
||||
}
|
||||
|
||||
export async function runSuite(args: string[]): Promise<number> {
|
||||
const ran = args.filter((a) => a !== "--no-build");
|
||||
const files = ran.length > 0 ? ran : [endToEnd];
|
||||
|
||||
if (!args.includes("--no-build")) {
|
||||
// Before the run, always. The artifacts are built from two other repositories, and a suite
|
||||
// that tests yesterday's binary reports on code nobody is looking at (novox/hq 04-ISSUES/005).
|
||||
const built = rebuild();
|
||||
if (built.length > 0) console.log(`built: ${built.join(", ")}\n`);
|
||||
}
|
||||
|
||||
const running = spawn(
|
||||
process.execPath,
|
||||
["--test", "--test-concurrency=1", "--experimental-strip-types",
|
||||
...files],
|
||||
{ stdio: ["inherit", "pipe", "inherit"] },
|
||||
);
|
||||
|
||||
let seen = "";
|
||||
running.stdout.on("data", (chunk: Buffer) => {
|
||||
// Passed through as it arrives: a suite that takes a quarter of an hour must not look hung.
|
||||
process.stdout.write(chunk);
|
||||
seen += chunk.toString();
|
||||
});
|
||||
|
||||
const code: number = await new Promise((resolve) => {
|
||||
running.on("close", (c) => resolve(c ?? 1));
|
||||
});
|
||||
|
||||
console.log("\n" + reportOn(counted(seen), (p, f) => record(p, f, files)));
|
||||
return code;
|
||||
}
|
||||
|
||||
/**
|
||||
* reportOn decides whether this run says anything worth recording, and records it if so.
|
||||
*
|
||||
* Separated from the spawning so the decision can be tested: **no receipt rather than a guessed
|
||||
* one** is the rule that keeps the record meaning something, and it was written where nothing
|
||||
* could check it — which is 04-ISSUES/005 in miniature, inside the fix for it.
|
||||
*/
|
||||
export function reportOn(
|
||||
counts: { passed: number | null; failed: number | null },
|
||||
write: (passed: number, failed: number) => { against: Record<string, string>; ran: string[] },
|
||||
): string {
|
||||
if (counts.passed === null || counts.failed === null) {
|
||||
// A run whose result could not be read is a run nobody can say anything about. Writing
|
||||
// "0 failed" because nothing said otherwise is how a green record comes to mean nothing.
|
||||
return "could not read what the runner reported; no receipt written";
|
||||
}
|
||||
const receipt = write(counts.passed, counts.failed);
|
||||
const against = Object.entries(receipt.against).map(([n, c]) => `${n} ${c}`).join(", ");
|
||||
return `recorded: ${receipt.ran.join(", ")} — ${counts.passed} passed, ` +
|
||||
`${counts.failed} failed, against ${against || "nothing in git"}`;
|
||||
}
|
||||
Reference in New Issue
Block a user