Files
hq/04-ISSUES/012-a-scenario-machine-is-too-small-for-four-images/00-report.md
T
jschoubben aafeb5c9df Three issues resolved: one closed by evidence, two answered by the replacement
012 named its own closing condition — a scenario with four images coming up —
and the scenario now stocks seven and has raised cleanly many times at the
memory the wrong diagnosis had raised.

001 is answered by the host reading the package database back after installing.
002 was NOT answered and was present here too, so it is a fix rather than a
note: a stale index is now named instead of reported as a failed install.
2026-08-31 13:00:03 +02:00

89 lines
4.4 KiB
Markdown

---
status: resolved
opened: 2026-08-30
located-in: [mesh-lab]
fixed-by: mesh-lab — scenario machines stay at 1 GiB, and the scenario now stocks seven images
amended-design:
---
# 012 — Raising a scenario machine's memory stops the substrate coming up
*Renamed after the first diagnosis turned out to be wrong. What that was, and how it was wrong, is
below — it is the more useful half of this report.*
## Symptom
Adding a fourth image to the two-machine scenario made the bootstrap fail every time. The store
container was created, and its readiness check then failed for the full three minutes with
**no output at all**:
```
failed store-ready (in mesh-store: ... pg_isready ... exit 1):
the action ran without error and its own verify still fails: docker exited 1:
```
Empty after the colon. The check runs `pg_isready` inside the container and prints the store's own
last lines when it gives up; producing nothing means **the container was not running**, which is a
different fault from a database that is slow to start.
Three images: the scenario raises, both machines join, and nine assertions pass. Four: it never
gets past the store. Reverting the fourth image restores it.
## The first diagnosis was wrong, and this is why it is worth writing down
**Two things changed at once.** A fourth image was added to the scenario, and — reasoning that a
machine running a database, a broker and the control plane at once is genuinely small — scenario
machines were raised from 1 GiB to 2 GiB. The bootstrap then failed every time, and the fourth
image was blamed.
**Removing the image did not fix it. Removing the memory increase did.** Nine assertions pass again
with the extra memory reverted, on a scenario with three images. So the cause is the memory change:
three machines at 2 GiB, on a host also running other work, contend enough that the store container
does not come up at all.
**The lesson is the ordinary one and it still caught me:** two changes went in together, the failure
was attributed to the plausible one, and an issue was written recording the wrong cause. What found
it was reverting to the exact last-known-good state rather than reverting the suspicious change.
**What remains untested** is whether a fourth image alone is fine. It probably is. Nothing has
measured it, and the honest state of this issue is that the thing it was opened about was never
demonstrated.
## What was worth keeping
**A better diagnostic.** The readiness check now prints what it saw before giving up. That is what
showed the output was empty, which is what said *the container is not running* rather than *the
database is slow* — and which will make the next occurrence of this a diagnosis instead of a
retry.
## What this blocks
The mesh running its **own artifact store** — a registry as a module — needs a registry image on
the machine so the module can mirror one, which is the fourth image. The module is written and its
manifest is accepted; what has not been proven is a machine assigned it serving artifacts to
another machine.
## How this will be checked
A scenario raised with four images comes up and passes the assertions that three do — which is
what was never actually established. Until then the artifact-store test is not in the shared
scenario, with a note saying where it went and why.
## Resolved
*2026-08-31.* The condition this report set was *a scenario raised with four images comes up and
passes the assertions that three do*. The scenario now stocks **seven** and has raised cleanly
many times over, with the machines at 1 GiB where the wrong diagnosis had put them at 2.
So both halves are settled. **The memory increase was the cause** — reverted, and never
reintroduced. **A fourth image was never the problem**, which this report said had not been
demonstrated either way, and now has been: three more were added on top of it, and the artifact
store the issue said was blocked is proven in the shared scenario rather than kept out of it.
**The diagnostic that came out of it is what remains valuable.** The readiness check prints what it
saw before giving up, which is what turned *the database is slow* into *the container is not
running*. It has since caught a different fault of the same shape — an action succeeding into a
state its own verify rejects
([04-ISSUES/017](../017-an-action-succeeded-into-a-state-its-verify-rejects/00-report.md)) — which
is the argument for keeping a good diagnostic after the incident that prompted it is gone.