024 — a run stalls before the host is placed, and says nothing while it does

Seen twice today. Once mid-run: thirteen passes, then the process ended
with no summary, no failure and no receipt. Once from the start: the
first test ran 35 minutes against a measured 4.5 and was still running
when it was stopped.

Ruled out rather than assumed: not memory (84 GiB free, no OOM), not the
daemon (the stalled machine answered `incus exec` immediately), and not
the changes under test — the anchor VM had no host log and no
containers, so the run never reached placing the host.

What changed just before is that the rebuild went from two artifacts to
six, and every one of them is pushed into the scenario's registry, which
is the step the second stall sat in. Recorded as what changed, not as
the diagnosis.

The reason this is an issue and not a slow test: the suite prints
nothing between starting a scenario and finishing its first test, so
four minutes and thirty-five look identical from outside, and the only
recourse is to guess. That is how a workstation was left unbootable in
August. And a run that ends silently after thirteen passes is a run
somebody may believe.
This commit is contained in:
2026-09-01 03:20:40 +02:00
parent 39916b26e9
commit 407416e6d0
@@ -0,0 +1,68 @@
---
status: open
opened: 2026-09-01
located-in: [mesh-lab]
fixed-by:
amended-design:
---
# 024 — A run stalls before the host is placed, and says nothing while it does
## Symptom
The first scenario of the end-to-end suite raises its machines and then stops. Observed twice on
2026-09-01, both times after the rebuild step grew:
- **Once mid-run**, during the rotation test: the suite had reported thirteen passes, then the
process ended with no summary, no failure and no receipt.
- **Once from the start**: `a bare machine becomes a mesh` ran for **35 minutes** against a
measured 4.5, produced no output at all, and was still running when it was stopped by hand.
Both machines were `RUNNING` throughout. The second one was interrogated directly: the anchor VM
answered, and had **no `mesh-host` log and no containers** — so the run had not reached placing
the host. It was stuck earlier, in stocking the scenario's registry.
## What is not the cause
- **Not memory.** 84 GiB available, no OOM in the kernel log.
- **Not the daemon.** `incus exec` into the stalled machine answered immediately.
- **Not the changes under test.** The credential work is applied after the host is placed, and the
host was never placed.
## What changed just before
The rebuild step went from two artifacts to six — the suite now builds every image the run uses,
rather than the control plane's alone
([`005`](../005-pipeline-test-harness-unbuildable/00-report.md)'s family, fixed the same day). More
images are built, and every one of them is then pushed into the scenario's own registry, which is
the step the second stall was sitting in.
That is a strong coincidence and not yet a diagnosis. It is written down as what changed, not as
the answer.
## Why it matters more than a slow test
**A stall is indistinguishable from work.** The suite prints nothing between starting a scenario
and finishing its first test, so four and a half minutes and thirty-five look identical from the
outside — and the operator's only recourse is to guess, which is precisely how a workstation was
left unbootable in August by killing a package manager that was working.
The first stall is worse: the process ended *silently* after thirteen passes. No summary, no
receipt, nothing that says the run was cut short. A run that stops without saying so is a run
somebody may believe.
## What a fix has to give
- **Progress while stocking**, so a long step is visibly a long step. Bytes moved, images pushed,
anything that changes.
- **A receipt when a run is cut short**, saying how far it got. `lastrun` already refuses to guess;
what is missing is it being written at all when the process dies mid-run.
- **A stated timeout on stocking**, so a stall ends as a failure with a reason rather than as a
process somebody eventually kills.
## What was done instead, for now
The run was stopped by hand and the instances removed. The gate that does not need a hypervisor —
`make check` in `mesh-control`, against a real PostgreSQL — is green, and the last complete run of
this suite was 21 of 24 with every failure diagnosed. **Neither of those covers what this suite
covers**, which is the point of the suite.