Three issues resolved: one closed by evidence, two answered by the replacement

012 named its own closing condition — a scenario with four images coming up —
and the scenario now stocks seven and has raised cleanly many times at the
memory the wrong diagnosis had raised.

001 is answered by the host reading the package database back after installing.
002 was NOT answered and was present here too, so it is a fix rather than a
note: a stale index is now named instead of reported as a failed install.
This commit is contained in:
2026-08-31 13:00:03 +02:00
parent 573a94e102
commit aafeb5c9df
3 changed files with 67 additions and 8 deletions
@@ -1,8 +1,8 @@
--- ---
status: open status: resolved
opened: 2026-08-22 opened: 2026-08-22
located-in: [] located-in: [mesh-host]
fixed-by: fixed-by: mesh-host — a package is read back from the package database after installing
amended-design: amended-design:
--- ---
@@ -47,3 +47,18 @@ This is an instance where it was never applied.
- Is this specific to package installation, or does the surrounding stage swallow every - Is this specific to package installation, or does the surrounding stage swallow every
non-zero exit? non-zero exit?
- The fix has been open for two days. What is the review path for a change of this class? - The fix has been open for two days. What is the review path for a change of this class?
## How it is answered
*2026-08-31.* **The host reads the package database back after installing**, and refuses when it
does not have the package:
> `<name> was installed without error and the package database does not have it`
That is the general rule this issue is one instance of, and the host applies it to everything it
does: a command exiting zero says a transaction was *accepted*, not that the machine changed. The
same read-back is why a container that starts and immediately dies fails an apply, and why a
service asked to run is checked rather than assumed.
HAL keeps the fault until its provisioning is switched off. Fixing it there would mean
implementing the read-back twice, in the system being replaced.
@@ -1,8 +1,8 @@
--- ---
status: open status: resolved
opened: 2026-08-22 opened: 2026-08-22
located-in: [] located-in: [mesh-host]
fixed-by: fixed-by: mesh-host — a stale package index is named rather than reported as a failed install
amended-design: amended-design:
--- ---
@@ -38,3 +38,29 @@ the job is green.
versions — or is an index sync part of the install step? versions — or is an index sync part of the install step?
- A partial sync is unsafe on the platform in use; a full upgrade is the only sanctioned fix. - A partial sync is unsafe on the platform in use; a full upgrade is the only sanctioned fix.
Does that make index freshness a scheduled node concern rather than a pipeline one? Does that make index freshness a scheduled node concern rather than a pipeline one?
## How it is answered
*2026-08-31. It was present in the replacement too, which is why this is a fix rather than a note
saying the new mesh does not have it.*
**The failure is named.** A machine asking for a version the mirrors have replaced now says so, and
says what fixes it — a full upgrade of the machine.
**It is deliberately not fixed by synchronising.** `pacman -Sy <pkg>` installs a package built
against libraries the machine does not have: a partial upgrade, unsupported on this distribution,
which surfaces much later as something apparently unrelated. That is a decision about the whole
machine, and a host that made it silently while applying one resource would be taking a large
decision in a small place.
So the host distinguishes the two cases and leaves the decision where it belongs. **A declaration
that is wrong and a machine that is out of date fail identically otherwise, and they are fixed in
completely different places.**
**And the package manager's own words were being thrown away** — the output was read into `_`, so
the 404s that name the cause never reached anybody. Whatever it said is now part of the failure,
which is the rule everywhere else here and was not being followed in the one place where the reason
exists only in the output.
*Checked by a stale-index failure being named as one, an ordinary missing package not being, a
single mirror timing out not being, and a successful install still saying nothing.*
@@ -1,8 +1,8 @@
--- ---
status: open status: resolved
opened: 2026-08-30 opened: 2026-08-30
located-in: [mesh-lab] located-in: [mesh-lab]
fixed-by: fixed-by: mesh-lab — scenario machines stay at 1 GiB, and the scenario now stocks seven images
amended-design: amended-design:
--- ---
@@ -68,3 +68,21 @@ another machine.
A scenario raised with four images comes up and passes the assertions that three do — which is A scenario raised with four images comes up and passes the assertions that three do — which is
what was never actually established. Until then the artifact-store test is not in the shared what was never actually established. Until then the artifact-store test is not in the shared
scenario, with a note saying where it went and why. scenario, with a note saying where it went and why.
## Resolved
*2026-08-31.* The condition this report set was *a scenario raised with four images comes up and
passes the assertions that three do*. The scenario now stocks **seven** and has raised cleanly
many times over, with the machines at 1 GiB where the wrong diagnosis had put them at 2.
So both halves are settled. **The memory increase was the cause** — reverted, and never
reintroduced. **A fourth image was never the problem**, which this report said had not been
demonstrated either way, and now has been: three more were added on top of it, and the artifact
store the issue said was blocked is proven in the shared scenario rather than kept out of it.
**The diagnostic that came out of it is what remains valuable.** The readiness check prints what it
saw before giving up, which is what turned *the database is slow* into *the container is not
running*. It has since caught a different fault of the same shape — an action succeeding into a
state its own verify rejects
([04-ISSUES/017](../017-an-action-succeeded-into-a-state-its-verify-rejects/00-report.md)) — which
is the argument for keeping a good diagnostic after the incident that prompted it is gone.