diff --git a/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md index ccef9d9..6dcd31f 100644 --- a/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md +++ b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md @@ -11,9 +11,9 @@ amended-design: ## What was observed -For at least seventeen minutes after the last operator action, the control-node's host applied all -327 of its resources every ~6.5 minutes without pause, replacing every container on the machine each -time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail, +For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's +host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on +the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail, and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine sat near 8. Nothing was converging: each pass ended and the next began four seconds later. @@ -50,7 +50,7 @@ if err != nil { A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does not know", but "the mesh states these names do not exist", to every machine at once. -## Why it cannot recover on its own +## Why it sustains itself The loop closes through the control plane's own database: @@ -66,9 +66,19 @@ The loop closes through the control plane's own database: mesh: reporting: nats: connection closed`. 5. Back to 1. -It is stable in its instability: every pass destroys the evidence the next pass needs to decide it -has nothing to do. Nothing outside the machine has to be wrong for this to continue, and nothing -about it stops. +Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing +outside the machine has to be wrong for it to continue. + +**It is metastable, not permanent.** It ran from 22:46 to 23:11 — five full replacements of every +container on the machine — and then stopped on its own, when one pass happened to read the store +during a window it was up, composed the same roster twice running, and found nothing to do. Load fell +from 7.8 to 1.7 and the machine returned to its five-minute idle tick. + +That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not +control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or +a larger store, would not have found it. An outage that clears itself after twenty-five minutes and +five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it +is the same fault, harder to catch. ## What it is not