Issue 152: the loop is metastable, and it cleared at 23:11

It ran 22:46-23:11, five full replacements, and stopped when a pass
happened to read the store during a window it was up. The record said
nothing about it stops; that was wrong. Exiting by luck is the finding,
not a mitigation.
This commit is contained in:
2026-09-29 23:27:51 +02:00
parent 3c2b4fc6b6
commit 18f37c25b2
@@ -11,9 +11,9 @@ amended-design:
## What was observed ## What was observed
For at least seventeen minutes after the last operator action, the control-node's host applied all For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's
327 of its resources every ~6.5 minutes without pause, replacing every container on the machine each host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on
time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail, the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail,
and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine
sat near 8. Nothing was converging: each pass ended and the next began four seconds later. sat near 8. Nothing was converging: each pass ended and the next began four seconds later.
@@ -50,7 +50,7 @@ if err != nil {
A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does
not know", but "the mesh states these names do not exist", to every machine at once. not know", but "the mesh states these names do not exist", to every machine at once.
## Why it cannot recover on its own ## Why it sustains itself
The loop closes through the control plane's own database: The loop closes through the control plane's own database:
@@ -66,9 +66,19 @@ The loop closes through the control plane's own database:
mesh: reporting: nats: connection closed`. mesh: reporting: nats: connection closed`.
5. Back to 1. 5. Back to 1.
It is stable in its instability: every pass destroys the evidence the next pass needs to decide it Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing
has nothing to do. Nothing outside the machine has to be wrong for this to continue, and nothing outside the machine has to be wrong for it to continue.
about it stops.
**It is metastable, not permanent.** It ran from 22:46 to 23:11 — five full replacements of every
container on the machine — and then stopped on its own, when one pass happened to read the store
during a window it was up, composed the same roster twice running, and found nothing to do. Load fell
from 7.8 to 1.7 and the machine returned to its five-minute idle tick.
That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not
control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or
a larger store, would not have found it. An outage that clears itself after twenty-five minutes and
five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it
is the same fault, harder to catch.
## What it is not ## What it is not