diff --git a/04-ISSUES/024-a-lab-run-stalls-before-the-host-is-placed/00-report.md b/04-ISSUES/024-a-lab-run-stalls-before-the-host-is-placed/00-report.md new file mode 100644 index 0000000..037910b --- /dev/null +++ b/04-ISSUES/024-a-lab-run-stalls-before-the-host-is-placed/00-report.md @@ -0,0 +1,68 @@ +--- +status: open +opened: 2026-09-01 +located-in: [mesh-lab] +fixed-by: +amended-design: +--- + +# 024 — A run stalls before the host is placed, and says nothing while it does + +## Symptom + +The first scenario of the end-to-end suite raises its machines and then stops. Observed twice on +2026-09-01, both times after the rebuild step grew: + +- **Once mid-run**, during the rotation test: the suite had reported thirteen passes, then the + process ended with no summary, no failure and no receipt. +- **Once from the start**: `a bare machine becomes a mesh` ran for **35 minutes** against a + measured 4.5, produced no output at all, and was still running when it was stopped by hand. + +Both machines were `RUNNING` throughout. The second one was interrogated directly: the anchor VM +answered, and had **no `mesh-host` log and no containers** — so the run had not reached placing +the host. It was stuck earlier, in stocking the scenario's registry. + +## What is not the cause + +- **Not memory.** 84 GiB available, no OOM in the kernel log. +- **Not the daemon.** `incus exec` into the stalled machine answered immediately. +- **Not the changes under test.** The credential work is applied after the host is placed, and the + host was never placed. + +## What changed just before + +The rebuild step went from two artifacts to six — the suite now builds every image the run uses, +rather than the control plane's alone +([`005`](../005-pipeline-test-harness-unbuildable/00-report.md)'s family, fixed the same day). More +images are built, and every one of them is then pushed into the scenario's own registry, which is +the step the second stall was sitting in. + +That is a strong coincidence and not yet a diagnosis. It is written down as what changed, not as +the answer. + +## Why it matters more than a slow test + +**A stall is indistinguishable from work.** The suite prints nothing between starting a scenario +and finishing its first test, so four and a half minutes and thirty-five look identical from the +outside — and the operator's only recourse is to guess, which is precisely how a workstation was +left unbootable in August by killing a package manager that was working. + +The first stall is worse: the process ended *silently* after thirteen passes. No summary, no +receipt, nothing that says the run was cut short. A run that stops without saying so is a run +somebody may believe. + +## What a fix has to give + +- **Progress while stocking**, so a long step is visibly a long step. Bytes moved, images pushed, + anything that changes. +- **A receipt when a run is cut short**, saying how far it got. `lastrun` already refuses to guess; + what is missing is it being written at all when the process dies mid-run. +- **A stated timeout on stocking**, so a stall ends as a failure with a reason rather than as a + process somebody eventually kills. + +## What was done instead, for now + +The run was stopped by hand and the instances removed. The gate that does not need a hypervisor — +`make check` in `mesh-control`, against a real PostgreSQL — is green, and the last complete run of +this suite was 21 of 24 with every failure diagnosed. **Neither of those covers what this suite +covers**, which is the point of the suite.