Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24
@@ -0,0 +1,68 @@
|
|||||||
|
---
|
||||||
|
status: open
|
||||||
|
opened: 2026-09-01
|
||||||
|
located-in: [mesh-lab]
|
||||||
|
fixed-by:
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 024 — A run stalls before the host is placed, and says nothing while it does
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
The first scenario of the end-to-end suite raises its machines and then stops. Observed twice on
|
||||||
|
2026-09-01, both times after the rebuild step grew:
|
||||||
|
|
||||||
|
- **Once mid-run**, during the rotation test: the suite had reported thirteen passes, then the
|
||||||
|
process ended with no summary, no failure and no receipt.
|
||||||
|
- **Once from the start**: `a bare machine becomes a mesh` ran for **35 minutes** against a
|
||||||
|
measured 4.5, produced no output at all, and was still running when it was stopped by hand.
|
||||||
|
|
||||||
|
Both machines were `RUNNING` throughout. The second one was interrogated directly: the anchor VM
|
||||||
|
answered, and had **no `mesh-host` log and no containers** — so the run had not reached placing
|
||||||
|
the host. It was stuck earlier, in stocking the scenario's registry.
|
||||||
|
|
||||||
|
## What is not the cause
|
||||||
|
|
||||||
|
- **Not memory.** 84 GiB available, no OOM in the kernel log.
|
||||||
|
- **Not the daemon.** `incus exec` into the stalled machine answered immediately.
|
||||||
|
- **Not the changes under test.** The credential work is applied after the host is placed, and the
|
||||||
|
host was never placed.
|
||||||
|
|
||||||
|
## What changed just before
|
||||||
|
|
||||||
|
The rebuild step went from two artifacts to six — the suite now builds every image the run uses,
|
||||||
|
rather than the control plane's alone
|
||||||
|
([`005`](../005-pipeline-test-harness-unbuildable/00-report.md)'s family, fixed the same day). More
|
||||||
|
images are built, and every one of them is then pushed into the scenario's own registry, which is
|
||||||
|
the step the second stall was sitting in.
|
||||||
|
|
||||||
|
That is a strong coincidence and not yet a diagnosis. It is written down as what changed, not as
|
||||||
|
the answer.
|
||||||
|
|
||||||
|
## Why it matters more than a slow test
|
||||||
|
|
||||||
|
**A stall is indistinguishable from work.** The suite prints nothing between starting a scenario
|
||||||
|
and finishing its first test, so four and a half minutes and thirty-five look identical from the
|
||||||
|
outside — and the operator's only recourse is to guess, which is precisely how a workstation was
|
||||||
|
left unbootable in August by killing a package manager that was working.
|
||||||
|
|
||||||
|
The first stall is worse: the process ended *silently* after thirteen passes. No summary, no
|
||||||
|
receipt, nothing that says the run was cut short. A run that stops without saying so is a run
|
||||||
|
somebody may believe.
|
||||||
|
|
||||||
|
## What a fix has to give
|
||||||
|
|
||||||
|
- **Progress while stocking**, so a long step is visibly a long step. Bytes moved, images pushed,
|
||||||
|
anything that changes.
|
||||||
|
- **A receipt when a run is cut short**, saying how far it got. `lastrun` already refuses to guess;
|
||||||
|
what is missing is it being written at all when the process dies mid-run.
|
||||||
|
- **A stated timeout on stocking**, so a stall ends as a failure with a reason rather than as a
|
||||||
|
process somebody eventually kills.
|
||||||
|
|
||||||
|
## What was done instead, for now
|
||||||
|
|
||||||
|
The run was stopped by hand and the instances removed. The gate that does not need a hypervisor —
|
||||||
|
`make check` in `mesh-control`, against a real PostgreSQL — is green, and the last complete run of
|
||||||
|
this suite was 21 of 24 with every failure diagnosed. **Neither of those covers what this suite
|
||||||
|
covers**, which is the point of the suite.
|
||||||
Reference in New Issue
Block a user