Say what the lab is doing, while it is doing it
novox/hq 04-ISSUES/024. A run stalled for thirty-five minutes and said nothing. The cause was a link systemd was still configuring, three layers down inside a `docker load` blocked on a socket — and every one of those layers knew what it was waiting for. None of them said so. Three decisions, each doing work. **Every external command is logged, at the three places that run one.** Ninety-seven call sites reach a hypervisor or a container runtime through three wrappers, so instrumenting the wrappers covers all of them and nothing has to remember to log. **A command still running says so while it runs.** A line before and a line after tells you nothing until the after arrives, which is exactly the case that matters. Anything outstanding past fifteen seconds reports itself with how long it has been going. It is reported as still running, not as stuck — which it is is not knowable from there, and a log that calls a slow step a hang teaches people to ignore it. **It goes to a file, written synchronously.** Node block-buffers stdout when redirected and a test runner buffers it again, so a console log can sit minutes behind. `appendFileSync` cannot lag. Two things this found in itself while being written, both the same shape as what it exists to catch: A question that answers no is not a fault. Half the lab's commands are questions — does this network exist, is the agent up yet — and they fail constantly while a scenario comes up. Logging those as faults filled a healthy run with ✗, which is how you end up ignoring ✗ when one is real. They are recorded quietly now, and still recorded. And `around` skipped its own wrapper when a step's level was below the configured one — taking the failure line and the heartbeat with it. The two things worth having at a low level were the two that vanished at exactly the level somebody would use. The gate belongs in `write`. Also unsilences the four call sites that passed a callback throwing everything away, including the one the stall sat in, and tees `raise`'s progress into the file whether or not a caller asked to see it — the end-to-end test passed no callback, so the one run that mattered reported not a single step.
This commit is contained in:
@@ -165,6 +165,27 @@ If it says the daemon is not reachable, the group grant postdates the shell. `ne
|
||||
but a heredoc into `newgrp` runs the suite as a child of a shell that then exits — start it with
|
||||
`setsid nohup … &` inside the heredoc, or the run dies with the shell that launched it.
|
||||
|
||||
## When something takes too long
|
||||
|
||||
```sh
|
||||
export MESH_LAB_LOG=info # or debug, or trace
|
||||
export MESH_LAB_LOG_FILE=/tmp/lab.log # unset writes to stderr
|
||||
```
|
||||
|
||||
| level | what it adds |
|
||||
|---|---|
|
||||
| `info` | each step of a raise, with how long the previous one took; anything that failed; **and a line every 15s naming whatever is still running** |
|
||||
| `debug` | every command the lab runs — incus, docker, and anything local — with its duration, and the stderr of anything that failed |
|
||||
| `trace` | what those commands printed |
|
||||
|
||||
**The heartbeat is the point.** A stall is a command that started and has not finished, and the
|
||||
only thing separating it from ordinary work is how long it has been going — which nothing can tell
|
||||
you unless something is still counting. At `info` a run says `… docker load -i /tmp/image.tar —
|
||||
still running after 45s` while it happens, rather than nothing until it gives up.
|
||||
|
||||
Written with `appendFileSync`, so unlike a redirected stdout it cannot lag behind the run. It is
|
||||
off unless asked for.
|
||||
|
||||
**A redirected log lags, so do not diagnose a stall from it.** Node block-buffers stdout when it
|
||||
is a file rather than a terminal, so `> run.log` can sit unchanged for minutes while the run is
|
||||
working normally. On 2026-09-01 that was read as a stall twice, once after a real stall had just
|
||||
|
||||
Reference in New Issue
Block a user