024 — a run stalls before the host is placed, and says nothing while it does
Seen twice today. Once mid-run: thirteen passes, then the process ended with no summary, no failure and no receipt. Once from the start: the first test ran 35 minutes against a measured 4.5 and was still running when it was stopped. Ruled out rather than assumed: not memory (84 GiB free, no OOM), not the daemon (the stalled machine answered `incus exec` immediately), and not the changes under test — the anchor VM had no host log and no containers, so the run never reached placing the host. What changed just before is that the rebuild went from two artifacts to six, and every one of them is pushed into the scenario's registry, which is the step the second stall sat in. Recorded as what changed, not as the diagnosis. The reason this is an issue and not a slow test: the suite prints nothing between starting a scenario and finishing its first test, so four minutes and thirty-five look identical from outside, and the only recourse is to guess. That is how a workstation was left unbootable in August. And a run that ends silently after thirteen passes is a run somebody may believe.
This commit is contained in:
@@ -0,0 +1,68 @@
|
|||||||
|
---
|
||||||
|
status: open
|
||||||
|
opened: 2026-09-01
|
||||||
|
located-in: [mesh-lab]
|
||||||
|
fixed-by:
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 024 — A run stalls before the host is placed, and says nothing while it does
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
The first scenario of the end-to-end suite raises its machines and then stops. Observed twice on
|
||||||
|
2026-09-01, both times after the rebuild step grew:
|
||||||
|
|
||||||
|
- **Once mid-run**, during the rotation test: the suite had reported thirteen passes, then the
|
||||||
|
process ended with no summary, no failure and no receipt.
|
||||||
|
- **Once from the start**: `a bare machine becomes a mesh` ran for **35 minutes** against a
|
||||||
|
measured 4.5, produced no output at all, and was still running when it was stopped by hand.
|
||||||
|
|
||||||
|
Both machines were `RUNNING` throughout. The second one was interrogated directly: the anchor VM
|
||||||
|
answered, and had **no `mesh-host` log and no containers** — so the run had not reached placing
|
||||||
|
the host. It was stuck earlier, in stocking the scenario's registry.
|
||||||
|
|
||||||
|
## What is not the cause
|
||||||
|
|
||||||
|
- **Not memory.** 84 GiB available, no OOM in the kernel log.
|
||||||
|
- **Not the daemon.** `incus exec` into the stalled machine answered immediately.
|
||||||
|
- **Not the changes under test.** The credential work is applied after the host is placed, and the
|
||||||
|
host was never placed.
|
||||||
|
|
||||||
|
## What changed just before
|
||||||
|
|
||||||
|
The rebuild step went from two artifacts to six — the suite now builds every image the run uses,
|
||||||
|
rather than the control plane's alone
|
||||||
|
([`005`](../005-pipeline-test-harness-unbuildable/00-report.md)'s family, fixed the same day). More
|
||||||
|
images are built, and every one of them is then pushed into the scenario's own registry, which is
|
||||||
|
the step the second stall was sitting in.
|
||||||
|
|
||||||
|
That is a strong coincidence and not yet a diagnosis. It is written down as what changed, not as
|
||||||
|
the answer.
|
||||||
|
|
||||||
|
## Why it matters more than a slow test
|
||||||
|
|
||||||
|
**A stall is indistinguishable from work.** The suite prints nothing between starting a scenario
|
||||||
|
and finishing its first test, so four and a half minutes and thirty-five look identical from the
|
||||||
|
outside — and the operator's only recourse is to guess, which is precisely how a workstation was
|
||||||
|
left unbootable in August by killing a package manager that was working.
|
||||||
|
|
||||||
|
The first stall is worse: the process ended *silently* after thirteen passes. No summary, no
|
||||||
|
receipt, nothing that says the run was cut short. A run that stops without saying so is a run
|
||||||
|
somebody may believe.
|
||||||
|
|
||||||
|
## What a fix has to give
|
||||||
|
|
||||||
|
- **Progress while stocking**, so a long step is visibly a long step. Bytes moved, images pushed,
|
||||||
|
anything that changes.
|
||||||
|
- **A receipt when a run is cut short**, saying how far it got. `lastrun` already refuses to guess;
|
||||||
|
what is missing is it being written at all when the process dies mid-run.
|
||||||
|
- **A stated timeout on stocking**, so a stall ends as a failure with a reason rather than as a
|
||||||
|
process somebody eventually kills.
|
||||||
|
|
||||||
|
## What was done instead, for now
|
||||||
|
|
||||||
|
The run was stopped by hand and the instances removed. The gate that does not need a hypervisor —
|
||||||
|
`make check` in `mesh-control`, against a real PostgreSQL — is green, and the last complete run of
|
||||||
|
this suite was 21 of 24 with every failure diagnosed. **Neither of those covers what this suite
|
||||||
|
covers**, which is the point of the suite.
|
||||||
Reference in New Issue
Block a user