Say what the lab is doing, while it is doing it

novox/hq 04-ISSUES/024. A run stalled for thirty-five minutes and said
nothing. The cause was a link systemd was still configuring, three
layers down inside a `docker load` blocked on a socket — and every one
of those layers knew what it was waiting for. None of them said so.

Three decisions, each doing work.

**Every external command is logged, at the three places that run one.**
Ninety-seven call sites reach a hypervisor or a container runtime
through three wrappers, so instrumenting the wrappers covers all of them
and nothing has to remember to log.

**A command still running says so while it runs.** A line before and a
line after tells you nothing until the after arrives, which is exactly
the case that matters. Anything outstanding past fifteen seconds reports
itself with how long it has been going. It is reported as still running,
not as stuck — which it is is not knowable from there, and a log that
calls a slow step a hang teaches people to ignore it.

**It goes to a file, written synchronously.** Node block-buffers stdout
when redirected and a test runner buffers it again, so a console log can
sit minutes behind. `appendFileSync` cannot lag.

Two things this found in itself while being written, both the same shape
as what it exists to catch:

A question that answers no is not a fault. Half the lab's commands are
questions — does this network exist, is the agent up yet — and they fail
constantly while a scenario comes up. Logging those as faults filled a
healthy run with ✗, which is how you end up ignoring ✗ when one is real.
They are recorded quietly now, and still recorded.

And `around` skipped its own wrapper when a step's level was below the
configured one — taking the failure line and the heartbeat with it. The
two things worth having at a low level were the two that vanished at
exactly the level somebody would use. The gate belongs in `write`.

Also unsilences the four call sites that passed a callback throwing
everything away, including the one the stall sat in, and tees `raise`'s
progress into the file whether or not a caller asked to see it — the
end-to-end test passed no callback, so the one run that mattered
reported not a single step.
This commit is contained in:
2026-09-01 10:43:38 +02:00
parent 5137720aa7
commit bb14ecb7e0
8 changed files with 464 additions and 19 deletions
+29 -1
View File
@@ -13,6 +13,8 @@
import { spawn } from "node:child_process";
import { around, log, shorten } from "../log.ts";
/**
* How to invoke incus. Overridable because the socket is group-owned and a session that
* predates the group grant cannot reach it — which is a real thing that happens on the
@@ -52,6 +54,27 @@ export class IncusError extends Error {
* that worked perfectly when typed.
*/
export async function incus(args: string[], timeoutMs = 60_000): Promise<IncusResult> {
// **Every command through here is recorded** (novox/hq 04-ISSUES/024). This is one of three
// places the lab runs an external program, and ninety-odd call sites reach a hypervisor through
// it — so logging here covers all of them and none of them has to remember to.
return invoke(args, timeoutMs, false);
}
/**
* `expectedToFail` is not about this command; it is about the caller.
*
* `incus` rejects and the caller is expected to care. `incusOk` and `succeeds` turn a failure into
* an answer — *does this network exist*, *is the agent up yet* — and those are asked constantly
* while a scenario comes up. Recording them as faults fills a healthy run with ✗.
*/
function invoke(args: string[], timeoutMs: number, expectedToFail: boolean): Promise<IncusResult> {
return around(`incus ${shorten(args)}`, () => run(args, timeoutMs), {
heartbeatMs: 15_000,
expectedToFail,
});
}
function run(args: string[], timeoutMs: number): Promise<IncusResult> {
const [command, ...prefix] = INCUS;
if (!command) throw new Error("MESH_LAB_INCUS is empty");
@@ -80,10 +103,15 @@ export async function incus(args: string[], timeoutMs = 60_000): Promise<IncusRe
child.on("close", (code) => {
clearTimeout(timer);
if (timedOut) {
// Said explicitly. A SIGKILL leaves an empty stderr, so without this the failure arrives
// with no explanation at all — which is how the first raise reported "(no output)".
log.info(`incus ${shorten(args)} was killed after ${timeoutMs}ms`);
reject(new IncusError(args, `timed out after ${timeoutMs}ms`, null));
} else if (code === 0) {
if (stdout.trim()) log.trace(` stdout: ${shorten([stdout.trim()], 400)}`);
resolve({ stdout, stderr });
} else {
log.debug(` exit ${code}: ${shorten([stderr.trim() || "(nothing on stderr)"], 400)}`);
reject(new IncusError(args, stderr, code));
}
});
@@ -105,7 +133,7 @@ export async function incus(args: string[], timeoutMs = 60_000): Promise<IncusRe
*/
export async function incusOk(args: string[], timeoutMs = 60_000): Promise<string | null> {
try {
return (await incus(args, timeoutMs)).stdout;
return (await invoke(args, timeoutMs, true)).stdout;
} catch {
return null;
}