Issue 012 — the first diagnosis was wrong, and that is the useful half
Two things changed at once: a fourth image in the scenario, and scenario machines raised from 1 GiB to 2 GiB. The bootstrap then failed every time, and the image was blamed. Removing the image did not fix it. Removing the memory increase did — nine assertions pass again on three images with the machines back at 1 GiB. Three machines at 2 GiB on a host doing other work contend enough that the store container does not come up at all. The ordinary lesson, and it still caught me: two changes together, the failure attributed to the plausible one, and an issue written recording the wrong cause. What found it was reverting to the exact last-known-good state rather than reverting the suspicious change. What remains untested is whether a fourth image alone is fine. Probably. Nothing has measured it, and the honest state of this issue is that what it was opened about was never demonstrated.
This commit is contained in:
@@ -6,7 +6,10 @@ fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 012 — A scenario machine cannot hold four images and raise a substrate
|
||||
# 012 — Raising a scenario machine's memory stops the substrate coming up
|
||||
|
||||
*Renamed after the first diagnosis turned out to be wrong. What that was, and how it was wrong, is
|
||||
below — it is the more useful half of this report.*
|
||||
|
||||
## Symptom
|
||||
|
||||
@@ -26,21 +29,32 @@ different fault from a database that is slow to start.
|
||||
Three images: the scenario raises, both machines join, and nine assertions pass. Four: it never
|
||||
gets past the store. Reverting the fourth image restores it.
|
||||
|
||||
## What was tried
|
||||
## The first diagnosis was wrong, and this is why it is worth writing down
|
||||
|
||||
- **More memory.** Scenario machines were raised from 1 GiB to 2 GiB, on the reasoning that a
|
||||
machine running a database, a broker and the control plane at once is genuinely small. It did not
|
||||
change the outcome, so memory is not it — the change is kept because the reasoning holds
|
||||
independently.
|
||||
- **A better diagnostic.** The readiness check now prints what it saw before giving up
|
||||
([04-ISSUES/011](../011-one-broken-module-blocks-every-other/00-report.md) changed the apply
|
||||
around it). That is what showed the output was empty, which is what rules out slowness.
|
||||
**Two things changed at once.** A fourth image was added to the scenario, and — reasoning that a
|
||||
machine running a database, a broker and the control plane at once is genuinely small — scenario
|
||||
machines were raised from 1 GiB to 2 GiB. The bootstrap then failed every time, and the fourth
|
||||
image was blamed.
|
||||
|
||||
## What is most likely, and untested
|
||||
**Removing the image did not fix it. Removing the memory increase did.** Nine assertions pass again
|
||||
with the extra memory reverted, on a scenario with three images. So the cause is the memory change:
|
||||
three machines at 2 GiB, on a host also running other work, contend enough that the store container
|
||||
does not come up at all.
|
||||
|
||||
**Disk.** Four images plus the base image on a machine whose root disk the scenario does not size.
|
||||
A database that cannot write its data directory does not start, and the container exits — which
|
||||
matches the evidence exactly. Nothing has measured it.
|
||||
**The lesson is the ordinary one and it still caught me:** two changes went in together, the failure
|
||||
was attributed to the plausible one, and an issue was written recording the wrong cause. What found
|
||||
it was reverting to the exact last-known-good state rather than reverting the suspicious change.
|
||||
|
||||
**What remains untested** is whether a fourth image alone is fine. It probably is. Nothing has
|
||||
measured it, and the honest state of this issue is that the thing it was opened about was never
|
||||
demonstrated.
|
||||
|
||||
## What was worth keeping
|
||||
|
||||
**A better diagnostic.** The readiness check now prints what it saw before giving up. That is what
|
||||
showed the output was empty, which is what said *the container is not running* rather than *the
|
||||
database is slow* — and which will make the next occurrence of this a diagnosis instead of a
|
||||
retry.
|
||||
|
||||
## What this blocks
|
||||
|
||||
@@ -51,5 +65,6 @@ another machine.
|
||||
|
||||
## How this will be checked
|
||||
|
||||
A scenario raised with four images comes up and passes the assertions that three do. Until then the
|
||||
artifact-store test is not in the shared scenario, with a note saying where it went and why.
|
||||
A scenario raised with four images comes up and passes the assertions that three do — which is
|
||||
what was never actually established. Until then the artifact-store test is not in the shared
|
||||
scenario, with a note saying where it went and why.
|
||||
|
||||
Reference in New Issue
Block a user