A run rebuilds what it tests, and leaves a receipt saying what it covered

The danger is not that the suite breaks. It is that nobody notices it
stopped running (novox/hq 04-ISSUES/005). The harness this replaces had
not built for two and a half months and nothing said so — and this suite
needs a hypervisor, so it inherits exactly that: it runs when somebody
remembers, and remembering is not a mechanism.

So running, recording, and rebuilding are one act:

- the host binary, control-plane image and builder are rebuilt from
  source first. The last two both parse manifests; building one and not
  the other left a binary eleven hours old refusing a field the mesh had
  just renamed, found by a full run.
- a receipt lands in XDG state — outside git, because the question is
  whether *this machine* has run it, and a receipt in git would be a
  claim about everybody's machine made by whoever committed last.
- `last-run` judges it and exits non-zero when it no longer counts.

Three faults found by running the thing rather than reading it, each now
held by a test confirmed to fail without it:

- counted() passed every test while parsing nothing. The runner colours
  its summary even into a pipe; the fixtures were clean text that had
  been imagined rather than captured. A fixture that agrees with the
  mistake proves the mistake.
- a receipt for `suite test/lastrun.test.ts` was indistinguishable from
  one for the real thing — 005's own symptom, rebuilt inside its remedy.
  The receipt now records what ran.
- a tree with uncommitted work reported the bare commit, claiming
  coverage of code nobody can check out. Nothing else could tell: the
  hash is identical either way.

Proven on real machines: 22/22, against all three repositories.
This commit is contained in:
2026-08-31 15:02:19 +02:00
parent e1317c9a69
commit 033ad7ec69
11 changed files with 762 additions and 15 deletions
+89
View File
@@ -0,0 +1,89 @@
/**
* Run the end-to-end suite, and leave a receipt saying it ran.
*
* **Here rather than inside the tests, because the tests cannot know their own totals.** Node's
* runner reports them to whatever invoked it, and a test file inventing its own count would be a
* receipt that says whatever the last edit made it say.
*
* Here rather than in a shell script for the same reason the rebuild is: a step that lives in
* somebody's terminal history is a step that runs when they remember (novox/hq 04-ISSUES/005).
*/
import { spawn } from "node:child_process";
import { endToEnd, record } from "./lastrun.ts";
import { rebuild } from "./rebuild.ts";
/** counted is what the runner said, or nulls when it said nothing recognisable. */
export function counted(output: string): { passed: number | null; failed: number | null } {
// The runner's own summary lines, each on a line of its own. Anchored, so a test *named*
// "pass 3" cannot be mistaken for the total — which is not a hypothetical worry in a suite whose
// tests are named in sentences.
// Stripped first: the runner colours its summary even when its stdout is a pipe, so the line is
// "\x1b[34m\u2139 pass 8\x1b[39m" and an anchored pattern never sees the start of it. Found by
// running this against the real runner — the fixture it was first written against was output I
// had imagined, which is a test that agrees with the mistake it was written beside.
const plain = output.replace(/\u001b\[[0-9;]*m/g, "");
const total = (what: RegExp) => {
const found = plain.match(what);
return found ? Number(found[1]) : null;
};
return {
passed: total(/^\s*(?:\u2139|#)\s*pass\s+(\d+)\s*$/m),
failed: total(/^\s*(?:\u2139|#)\s*fail\s+(\d+)\s*$/m),
};
}
export async function runSuite(args: string[]): Promise<number> {
const ran = args.filter((a) => a !== "--no-build");
const files = ran.length > 0 ? ran : [endToEnd];
if (!args.includes("--no-build")) {
// Before the run, always. The artifacts are built from two other repositories, and a suite
// that tests yesterday's binary reports on code nobody is looking at (novox/hq 04-ISSUES/005).
const built = rebuild();
if (built.length > 0) console.log(`built: ${built.join(", ")}\n`);
}
const running = spawn(
process.execPath,
["--test", "--test-concurrency=1", "--experimental-strip-types",
...files],
{ stdio: ["inherit", "pipe", "inherit"] },
);
let seen = "";
running.stdout.on("data", (chunk: Buffer) => {
// Passed through as it arrives: a suite that takes a quarter of an hour must not look hung.
process.stdout.write(chunk);
seen += chunk.toString();
});
const code: number = await new Promise((resolve) => {
running.on("close", (c) => resolve(c ?? 1));
});
console.log("\n" + reportOn(counted(seen), (p, f) => record(p, f, files)));
return code;
}
/**
* reportOn decides whether this run says anything worth recording, and records it if so.
*
* Separated from the spawning so the decision can be tested: **no receipt rather than a guessed
* one** is the rule that keeps the record meaning something, and it was written where nothing
* could check it — which is 04-ISSUES/005 in miniature, inside the fix for it.
*/
export function reportOn(
counts: { passed: number | null; failed: number | null },
write: (passed: number, failed: number) => { against: Record<string, string>; ran: string[] },
): string {
if (counts.passed === null || counts.failed === null) {
// A run whose result could not be read is a run nobody can say anything about. Writing
// "0 failed" because nothing said otherwise is how a green record comes to mean nothing.
return "could not read what the runner reported; no receipt written";
}
const receipt = write(counts.passed, counts.failed);
const against = Object.entries(receipt.against).map(([n, c]) => `${n} ${c}`).join(", ");
return `recorded: ${receipt.ran.join(", ")} — ${counts.passed} passed, ` +
`${counts.failed} failed, against ${against || "nothing in git"}`;
}