Issue 152: the loop is metastable, and it cleared at 23:11

It ran 22:46-23:11, five full replacements, and stopped when a pass
happened to read the store during a window it was up. The record said
nothing about it stops; that was wrong. Exiting by luck is the finding,
not a mitigation.
This commit is contained in:
2026-09-29 23:27:51 +02:00
parent 3c2b4fc6b6
commit 18f37c25b2
@@ -11,9 +11,9 @@ amended-design:
## What was observed
For at least seventeen minutes after the last operator action, the control-node's host applied all
327 of its resources every ~6.5 minutes without pause, replacing every container on the machine each
time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail,
For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's
host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on
the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail,
and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine
sat near 8. Nothing was converging: each pass ended and the next began four seconds later.
@@ -50,7 +50,7 @@ if err != nil {
A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does
not know", but "the mesh states these names do not exist", to every machine at once.
## Why it cannot recover on its own
## Why it sustains itself
The loop closes through the control plane's own database:
@@ -66,9 +66,19 @@ The loop closes through the control plane's own database:
mesh: reporting: nats: connection closed`.
5. Back to 1.
It is stable in its instability: every pass destroys the evidence the next pass needs to decide it
has nothing to do. Nothing outside the machine has to be wrong for this to continue, and nothing
about it stops.
Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing
outside the machine has to be wrong for it to continue.
**It is metastable, not permanent.** It ran from 22:46 to 23:11 — five full replacements of every
container on the machine — and then stopped on its own, when one pass happened to read the store
during a window it was up, composed the same roster twice running, and found nothing to do. Load fell
from 7.8 to 1.7 and the machine returned to its five-minute idle tick.
That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not
control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or
a larger store, would not have found it. An outage that clears itself after twenty-five minutes and
five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it
is the same fault, harder to catch.
## What it is not