From aafeb5c9df7f44de553364f0c3ade25f450a9d27 Mon Sep 17 00:00:00 2001 From: jochen Date: Mon, 31 Aug 2026 13:00:03 +0200 Subject: [PATCH] Three issues resolved: one closed by evidence, two answered by the replacement MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 012 named its own closing condition — a scenario with four images coming up — and the scenario now stocks seven and has raised cleanly many times at the memory the wrong diagnosis had raised. 001 is answered by the host reading the package database back after installing. 002 was NOT answered and was present here too, so it is a fix rather than a note: a stale index is now named instead of reported as a failed install. --- .../00-report.md | 21 ++++++++++-- .../00-report.md | 32 +++++++++++++++++-- .../00-report.md | 22 +++++++++++-- 3 files changed, 67 insertions(+), 8 deletions(-) diff --git a/04-ISSUES/001-failed-package-install-reports-success/00-report.md b/04-ISSUES/001-failed-package-install-reports-success/00-report.md index a45888a..0b61223 100644 --- a/04-ISSUES/001-failed-package-install-reports-success/00-report.md +++ b/04-ISSUES/001-failed-package-install-reports-success/00-report.md @@ -1,8 +1,8 @@ --- -status: open +status: resolved opened: 2026-08-22 -located-in: [] -fixed-by: +located-in: [mesh-host] +fixed-by: mesh-host — a package is read back from the package database after installing amended-design: --- @@ -47,3 +47,18 @@ This is an instance where it was never applied. - Is this specific to package installation, or does the surrounding stage swallow every non-zero exit? - The fix has been open for two days. What is the review path for a change of this class? + +## How it is answered + +*2026-08-31.* **The host reads the package database back after installing**, and refuses when it +does not have the package: + +> ` was installed without error and the package database does not have it` + +That is the general rule this issue is one instance of, and the host applies it to everything it +does: a command exiting zero says a transaction was *accepted*, not that the machine changed. The +same read-back is why a container that starts and immediately dies fails an apply, and why a +service asked to run is checked rather than assumed. + +HAL keeps the fault until its provisioning is switched off. Fixing it there would mean +implementing the read-back twice, in the system being replaced. diff --git a/04-ISSUES/002-stale-package-index-fails-silently/00-report.md b/04-ISSUES/002-stale-package-index-fails-silently/00-report.md index ccbc7b7..a9bfb5e 100644 --- a/04-ISSUES/002-stale-package-index-fails-silently/00-report.md +++ b/04-ISSUES/002-stale-package-index-fails-silently/00-report.md @@ -1,8 +1,8 @@ --- -status: open +status: resolved opened: 2026-08-22 -located-in: [] -fixed-by: +located-in: [mesh-host] +fixed-by: mesh-host — a stale package index is named rather than reported as a failed install amended-design: --- @@ -38,3 +38,29 @@ the job is green. versions — or is an index sync part of the install step? - A partial sync is unsafe on the platform in use; a full upgrade is the only sanctioned fix. Does that make index freshness a scheduled node concern rather than a pipeline one? + +## How it is answered + +*2026-08-31. It was present in the replacement too, which is why this is a fix rather than a note +saying the new mesh does not have it.* + +**The failure is named.** A machine asking for a version the mirrors have replaced now says so, and +says what fixes it — a full upgrade of the machine. + +**It is deliberately not fixed by synchronising.** `pacman -Sy ` installs a package built +against libraries the machine does not have: a partial upgrade, unsupported on this distribution, +which surfaces much later as something apparently unrelated. That is a decision about the whole +machine, and a host that made it silently while applying one resource would be taking a large +decision in a small place. + +So the host distinguishes the two cases and leaves the decision where it belongs. **A declaration +that is wrong and a machine that is out of date fail identically otherwise, and they are fixed in +completely different places.** + +**And the package manager's own words were being thrown away** — the output was read into `_`, so +the 404s that name the cause never reached anybody. Whatever it said is now part of the failure, +which is the rule everywhere else here and was not being followed in the one place where the reason +exists only in the output. + +*Checked by a stale-index failure being named as one, an ordinary missing package not being, a +single mirror timing out not being, and a successful install still saying nothing.* diff --git a/04-ISSUES/012-a-scenario-machine-is-too-small-for-four-images/00-report.md b/04-ISSUES/012-a-scenario-machine-is-too-small-for-four-images/00-report.md index 7559c17..3b93ba6 100644 --- a/04-ISSUES/012-a-scenario-machine-is-too-small-for-four-images/00-report.md +++ b/04-ISSUES/012-a-scenario-machine-is-too-small-for-four-images/00-report.md @@ -1,8 +1,8 @@ --- -status: open +status: resolved opened: 2026-08-30 located-in: [mesh-lab] -fixed-by: +fixed-by: mesh-lab — scenario machines stay at 1 GiB, and the scenario now stocks seven images amended-design: --- @@ -68,3 +68,21 @@ another machine. A scenario raised with four images comes up and passes the assertions that three do — which is what was never actually established. Until then the artifact-store test is not in the shared scenario, with a note saying where it went and why. + +## Resolved + +*2026-08-31.* The condition this report set was *a scenario raised with four images comes up and +passes the assertions that three do*. The scenario now stocks **seven** and has raised cleanly +many times over, with the machines at 1 GiB where the wrong diagnosis had put them at 2. + +So both halves are settled. **The memory increase was the cause** — reverted, and never +reintroduced. **A fourth image was never the problem**, which this report said had not been +demonstrated either way, and now has been: three more were added on top of it, and the artifact +store the issue said was blocked is proven in the shared scenario rather than kept out of it. + +**The diagnostic that came out of it is what remains valuable.** The readiness check prints what it +saw before giving up, which is what turned *the database is slow* into *the container is not +running*. It has since caught a different fault of the same shape — an action succeeding into a +state its own verify rejects +([04-ISSUES/017](../017-an-action-succeeded-into-a-state-its-verify-rejects/00-report.md)) — which +is the argument for keeping a good diagnostic after the incident that prompted it is gone.