Say what the lab is doing, while it is doing it

novox/hq 04-ISSUES/024. A run stalled for thirty-five minutes and said
nothing. The cause was a link systemd was still configuring, three
layers down inside a `docker load` blocked on a socket — and every one
of those layers knew what it was waiting for. None of them said so.

Three decisions, each doing work.

**Every external command is logged, at the three places that run one.**
Ninety-seven call sites reach a hypervisor or a container runtime
through three wrappers, so instrumenting the wrappers covers all of them
and nothing has to remember to log.

**A command still running says so while it runs.** A line before and a
line after tells you nothing until the after arrives, which is exactly
the case that matters. Anything outstanding past fifteen seconds reports
itself with how long it has been going. It is reported as still running,
not as stuck — which it is is not knowable from there, and a log that
calls a slow step a hang teaches people to ignore it.

**It goes to a file, written synchronously.** Node block-buffers stdout
when redirected and a test runner buffers it again, so a console log can
sit minutes behind. `appendFileSync` cannot lag.

Two things this found in itself while being written, both the same shape
as what it exists to catch:

A question that answers no is not a fault. Half the lab's commands are
questions — does this network exist, is the agent up yet — and they fail
constantly while a scenario comes up. Logging those as faults filled a
healthy run with ✗, which is how you end up ignoring ✗ when one is real.
They are recorded quietly now, and still recorded.

And `around` skipped its own wrapper when a step's level was below the
configured one — taking the failure line and the heartbeat with it. The
two things worth having at a low level were the two that vanished at
exactly the level somebody would use. The gate belongs in `write`.

Also unsilences the four call sites that passed a callback throwing
everything away, including the one the stall sat in, and tees `raise`'s
progress into the file whether or not a caller asked to see it — the
end-to-end test passed no callback, so the one run that mattered
reported not a single step.
This commit is contained in:
2026-09-01 10:43:38 +02:00
parent 5137720aa7
commit bb14ecb7e0
8 changed files with 464 additions and 19 deletions
+21
View File
@@ -165,6 +165,27 @@ If it says the daemon is not reachable, the group grant postdates the shell. `ne
but a heredoc into `newgrp` runs the suite as a child of a shell that then exits — start it with
`setsid nohup … &` inside the heredoc, or the run dies with the shell that launched it.
## When something takes too long
```sh
export MESH_LAB_LOG=info # or debug, or trace
export MESH_LAB_LOG_FILE=/tmp/lab.log # unset writes to stderr
```
| level | what it adds |
|---|---|
| `info` | each step of a raise, with how long the previous one took; anything that failed; **and a line every 15s naming whatever is still running** |
| `debug` | every command the lab runs — incus, docker, and anything local — with its duration, and the stderr of anything that failed |
| `trace` | what those commands printed |
**The heartbeat is the point.** A stall is a command that started and has not finished, and the
only thing separating it from ordinary work is how long it has been going — which nothing can tell
you unless something is still counting. At `info` a run says `… docker load -i /tmp/image.tar —
still running after 45s` while it happens, rather than nothing until it gives up.
Written with `appendFileSync`, so unlike a redirected stdout it cannot lag behind the run. It is
off unless asked for.
**A redirected log lags, so do not diagnose a stall from it.** Node block-buffers stdout when it
is a file rather than a terminal, so `> run.log` can sit unchanged for minutes while the run is
working normally. On 2026-09-01 that was read as a stall twice, once after a real stall had just