Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24
@@ -0,0 +1,68 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-01
|
||||
located-in: [mesh-lab]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 024 — A run stalls before the host is placed, and says nothing while it does
|
||||
|
||||
## Symptom
|
||||
|
||||
The first scenario of the end-to-end suite raises its machines and then stops. Observed twice on
|
||||
2026-09-01, both times after the rebuild step grew:
|
||||
|
||||
- **Once mid-run**, during the rotation test: the suite had reported thirteen passes, then the
|
||||
process ended with no summary, no failure and no receipt.
|
||||
- **Once from the start**: `a bare machine becomes a mesh` ran for **35 minutes** against a
|
||||
measured 4.5, produced no output at all, and was still running when it was stopped by hand.
|
||||
|
||||
Both machines were `RUNNING` throughout. The second one was interrogated directly: the anchor VM
|
||||
answered, and had **no `mesh-host` log and no containers** — so the run had not reached placing
|
||||
the host. It was stuck earlier, in stocking the scenario's registry.
|
||||
|
||||
## What is not the cause
|
||||
|
||||
- **Not memory.** 84 GiB available, no OOM in the kernel log.
|
||||
- **Not the daemon.** `incus exec` into the stalled machine answered immediately.
|
||||
- **Not the changes under test.** The credential work is applied after the host is placed, and the
|
||||
host was never placed.
|
||||
|
||||
## What changed just before
|
||||
|
||||
The rebuild step went from two artifacts to six — the suite now builds every image the run uses,
|
||||
rather than the control plane's alone
|
||||
([`005`](../005-pipeline-test-harness-unbuildable/00-report.md)'s family, fixed the same day). More
|
||||
images are built, and every one of them is then pushed into the scenario's own registry, which is
|
||||
the step the second stall was sitting in.
|
||||
|
||||
That is a strong coincidence and not yet a diagnosis. It is written down as what changed, not as
|
||||
the answer.
|
||||
|
||||
## Why it matters more than a slow test
|
||||
|
||||
**A stall is indistinguishable from work.** The suite prints nothing between starting a scenario
|
||||
and finishing its first test, so four and a half minutes and thirty-five look identical from the
|
||||
outside — and the operator's only recourse is to guess, which is precisely how a workstation was
|
||||
left unbootable in August by killing a package manager that was working.
|
||||
|
||||
The first stall is worse: the process ended *silently* after thirteen passes. No summary, no
|
||||
receipt, nothing that says the run was cut short. A run that stops without saying so is a run
|
||||
somebody may believe.
|
||||
|
||||
## What a fix has to give
|
||||
|
||||
- **Progress while stocking**, so a long step is visibly a long step. Bytes moved, images pushed,
|
||||
anything that changes.
|
||||
- **A receipt when a run is cut short**, saying how far it got. `lastrun` already refuses to guess;
|
||||
what is missing is it being written at all when the process dies mid-run.
|
||||
- **A stated timeout on stocking**, so a stall ends as a failure with a reason rather than as a
|
||||
process somebody eventually kills.
|
||||
|
||||
## What was done instead, for now
|
||||
|
||||
The run was stopped by hand and the instances removed. The gate that does not need a hypervisor —
|
||||
`make check` in `mesh-control`, against a real PostgreSQL — is green, and the last complete run of
|
||||
this suite was 21 of 24 with every failure diagnosed. **Neither of those covers what this suite
|
||||
covers**, which is the point of the suite.
|
||||
Reference in New Issue
Block a user