Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24

Merged
jschoubben merged 177 commits from reconcile-init-into-main into main 2026-09-05 10:27:11 +00:00
Showing only changes of commit 407416e6d0 - Show all commits
@@ -0,0 +1,68 @@
---
status: open
opened: 2026-09-01
located-in: [mesh-lab]
fixed-by:
amended-design:
---
# 024 — A run stalls before the host is placed, and says nothing while it does
## Symptom
The first scenario of the end-to-end suite raises its machines and then stops. Observed twice on
2026-09-01, both times after the rebuild step grew:
- **Once mid-run**, during the rotation test: the suite had reported thirteen passes, then the
process ended with no summary, no failure and no receipt.
- **Once from the start**: `a bare machine becomes a mesh` ran for **35 minutes** against a
measured 4.5, produced no output at all, and was still running when it was stopped by hand.
Both machines were `RUNNING` throughout. The second one was interrogated directly: the anchor VM
answered, and had **no `mesh-host` log and no containers** — so the run had not reached placing
the host. It was stuck earlier, in stocking the scenario's registry.
## What is not the cause
- **Not memory.** 84 GiB available, no OOM in the kernel log.
- **Not the daemon.** `incus exec` into the stalled machine answered immediately.
- **Not the changes under test.** The credential work is applied after the host is placed, and the
host was never placed.
## What changed just before
The rebuild step went from two artifacts to six — the suite now builds every image the run uses,
rather than the control plane's alone
([`005`](../005-pipeline-test-harness-unbuildable/00-report.md)'s family, fixed the same day). More
images are built, and every one of them is then pushed into the scenario's own registry, which is
the step the second stall was sitting in.
That is a strong coincidence and not yet a diagnosis. It is written down as what changed, not as
the answer.
## Why it matters more than a slow test
**A stall is indistinguishable from work.** The suite prints nothing between starting a scenario
and finishing its first test, so four and a half minutes and thirty-five look identical from the
outside — and the operator's only recourse is to guess, which is precisely how a workstation was
left unbootable in August by killing a package manager that was working.
The first stall is worse: the process ended *silently* after thirteen passes. No summary, no
receipt, nothing that says the run was cut short. A run that stops without saying so is a run
somebody may believe.
## What a fix has to give
- **Progress while stocking**, so a long step is visibly a long step. Bytes moved, images pushed,
anything that changes.
- **A receipt when a run is cut short**, saying how far it got. `lastrun` already refuses to guess;
what is missing is it being written at all when the process dies mid-run.
- **A stated timeout on stocking**, so a stall ends as a failure with a reason rather than as a
process somebody eventually kills.
## What was done instead, for now
The run was stopped by hand and the instances removed. The gate that does not need a hypervisor —
`make check` in `mesh-control`, against a real PostgreSQL — is green, and the last complete run of
this suite was 21 of 24 with every failure diagnosed. **Neither of those covers what this suite
covers**, which is the point of the suite.