--- status: open opened: 2026-09-01 located-in: [mesh-lab] fixed-by: amended-design: --- # 024 — A run stalls before the host is placed, and says nothing while it does ## Symptom The first scenario of the end-to-end suite raises its machines and then stops. Observed twice on 2026-09-01, both times after the rebuild step grew: - **Once mid-run**, during the rotation test: the suite had reported thirteen passes, then the process ended with no summary, no failure and no receipt. - **Once from the start**: `a bare machine becomes a mesh` ran for **35 minutes** against a measured 4.5, produced no output at all, and was still running when it was stopped by hand. Both machines were `RUNNING` throughout. The second one was interrogated directly: the anchor VM answered, and had **no `mesh-host` log and no containers** — so the run had not reached placing the host. It was stuck earlier, in stocking the scenario's registry. ## What is not the cause - **Not memory.** 84 GiB available, no OOM in the kernel log. - **Not the daemon.** `incus exec` into the stalled machine answered immediately. - **Not the changes under test.** The credential work is applied after the host is placed, and the host was never placed. ## What changed just before The rebuild step went from two artifacts to six — the suite now builds every image the run uses, rather than the control plane's alone ([`005`](../005-pipeline-test-harness-unbuildable/00-report.md)'s family, fixed the same day). More images are built, and every one of them is then pushed into the scenario's own registry, which is the step the second stall was sitting in. That is a strong coincidence and not yet a diagnosis. It is written down as what changed, not as the answer. ## Why it matters more than a slow test **A stall is indistinguishable from work.** The suite prints nothing between starting a scenario and finishing its first test, so four and a half minutes and thirty-five look identical from the outside — and the operator's only recourse is to guess, which is precisely how a workstation was left unbootable in August by killing a package manager that was working. The first stall is worse: the process ended *silently* after thirteen passes. No summary, no receipt, nothing that says the run was cut short. A run that stops without saying so is a run somebody may believe. ## What a fix has to give - **Progress while stocking**, so a long step is visibly a long step. Bytes moved, images pushed, anything that changes. - **A receipt when a run is cut short**, saying how far it got. `lastrun` already refuses to guess; what is missing is it being written at all when the process dies mid-run. - **A stated timeout on stocking**, so a stall ends as a failure with a reason rather than as a process somebody eventually kills. ## What was done instead, for now The run was stopped by hand and the instances removed. The gate that does not need a hypervisor — `make check` in `mesh-control`, against a real PostgreSQL — is green, and the last complete run of this suite was 21 of 24 with every failure diagnosed. **Neither of those covers what this suite covers**, which is the point of the suite.