Files
hq/04-ISSUES/001-failed-package-install-reports-success/00-report.md
jschoubben aafeb5c9df Three issues resolved: one closed by evidence, two answered by the replacement
012 named its own closing condition — a scenario with four images coming up —
and the scenario now stocks seven and has raised cleanly many times at the
memory the wrong diagnosis had raised.

001 is answered by the host reading the package database back after installing.
002 was NOT answered and was present here too, so it is a fix rather than a
note: a stale index is now named instead of reported as a failed install.
2026-08-31 13:00:03 +02:00

2.3 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
resolved 2026-08-22
mesh-host
mesh-host — a package is read back from the package database after installing

001 — A failed package install does not fail the job

Symptom

A module declared a package. The install produced, from every mirror:

error: failed retrieving file … 404

followed by:

-> error installing repo packages

The prepare job then reported success. The package is absent; the pipeline is green.

Why this matters more than one missing package

The first thing the lab work asked the mesh to install demonstrated the exact fault the lab exists to catch — a step that failed, reported success, and left the next step to run against state that was never produced.

It is also a direct violation of a decision already taken and recorded: ADR 0010 says a step that fails must fail the job. That record notes the rule is applied instance by instance and enforced by no mechanism. This is an instance where it was never applied.

Evidence

  • Observed 2026-08-22 while declaring the virtualisation package required by ADR 0016.
  • A fix is written and open as a pull request, unmerged since 2026-08-20.

Open questions

  • Why is the failure swallowed — is the exit status discarded, or never checked?
  • Is this specific to package installation, or does the surrounding stage swallow every non-zero exit?
  • The fix has been open for two days. What is the review path for a change of this class?

How it is answered

2026-08-31. The host reads the package database back after installing, and refuses when it does not have the package:

<name> was installed without error and the package database does not have it

That is the general rule this issue is one instance of, and the host applies it to everything it does: a command exiting zero says a transaction was accepted, not that the machine changed. The same read-back is why a container that starts and immediately dies fails an apply, and why a service asked to run is checked rather than assumed.

HAL keeps the fault until its provisioning is switched off. Fixing it there would mean implementing the read-back twice, in the system being replaced.