From ba14e2b6295878705edecda930c587abb3699223 Mon Sep 17 00:00:00 2001 From: jochen Date: Sun, 30 Aug 2026 20:40:54 +0200 Subject: [PATCH] =?UTF-8?q?Issue=20012=20=E2=80=94=20the=20first=20diagnos?= =?UTF-8?q?is=20was=20wrong,=20and=20that=20is=20the=20useful=20half?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two things changed at once: a fourth image in the scenario, and scenario machines raised from 1 GiB to 2 GiB. The bootstrap then failed every time, and the image was blamed. Removing the image did not fix it. Removing the memory increase did — nine assertions pass again on three images with the machines back at 1 GiB. Three machines at 2 GiB on a host doing other work contend enough that the store container does not come up at all. The ordinary lesson, and it still caught me: two changes together, the failure attributed to the plausible one, and an issue written recording the wrong cause. What found it was reverting to the exact last-known-good state rather than reverting the suspicious change. What remains untested is whether a fourth image alone is fine. Probably. Nothing has measured it, and the honest state of this issue is that what it was opened about was never demonstrated. --- .../00-report.md | 45 ++++++++++++------- 1 file changed, 30 insertions(+), 15 deletions(-) diff --git a/04-ISSUES/012-a-scenario-machine-is-too-small-for-four-images/00-report.md b/04-ISSUES/012-a-scenario-machine-is-too-small-for-four-images/00-report.md index 1fbe53a..7559c17 100644 --- a/04-ISSUES/012-a-scenario-machine-is-too-small-for-four-images/00-report.md +++ b/04-ISSUES/012-a-scenario-machine-is-too-small-for-four-images/00-report.md @@ -6,7 +6,10 @@ fixed-by: amended-design: --- -# 012 — A scenario machine cannot hold four images and raise a substrate +# 012 — Raising a scenario machine's memory stops the substrate coming up + +*Renamed after the first diagnosis turned out to be wrong. What that was, and how it was wrong, is +below — it is the more useful half of this report.* ## Symptom @@ -26,21 +29,32 @@ different fault from a database that is slow to start. Three images: the scenario raises, both machines join, and nine assertions pass. Four: it never gets past the store. Reverting the fourth image restores it. -## What was tried +## The first diagnosis was wrong, and this is why it is worth writing down -- **More memory.** Scenario machines were raised from 1 GiB to 2 GiB, on the reasoning that a - machine running a database, a broker and the control plane at once is genuinely small. It did not - change the outcome, so memory is not it — the change is kept because the reasoning holds - independently. -- **A better diagnostic.** The readiness check now prints what it saw before giving up - ([04-ISSUES/011](../011-one-broken-module-blocks-every-other/00-report.md) changed the apply - around it). That is what showed the output was empty, which is what rules out slowness. +**Two things changed at once.** A fourth image was added to the scenario, and — reasoning that a +machine running a database, a broker and the control plane at once is genuinely small — scenario +machines were raised from 1 GiB to 2 GiB. The bootstrap then failed every time, and the fourth +image was blamed. -## What is most likely, and untested +**Removing the image did not fix it. Removing the memory increase did.** Nine assertions pass again +with the extra memory reverted, on a scenario with three images. So the cause is the memory change: +three machines at 2 GiB, on a host also running other work, contend enough that the store container +does not come up at all. -**Disk.** Four images plus the base image on a machine whose root disk the scenario does not size. -A database that cannot write its data directory does not start, and the container exits — which -matches the evidence exactly. Nothing has measured it. +**The lesson is the ordinary one and it still caught me:** two changes went in together, the failure +was attributed to the plausible one, and an issue was written recording the wrong cause. What found +it was reverting to the exact last-known-good state rather than reverting the suspicious change. + +**What remains untested** is whether a fourth image alone is fine. It probably is. Nothing has +measured it, and the honest state of this issue is that the thing it was opened about was never +demonstrated. + +## What was worth keeping + +**A better diagnostic.** The readiness check now prints what it saw before giving up. That is what +showed the output was empty, which is what said *the container is not running* rather than *the +database is slow* — and which will make the next occurrence of this a diagnosis instead of a +retry. ## What this blocks @@ -51,5 +65,6 @@ another machine. ## How this will be checked -A scenario raised with four images comes up and passes the assertions that three do. Until then the -artifact-store test is not in the shared scenario, with a note saying where it went and why. +A scenario raised with four images comes up and passes the assertions that three do — which is +what was never actually established. Until then the artifact-store test is not in the shared +scenario, with a note saying where it went and why.