From 18f37c25b2ef0cba6d56d51ebe697db6f9aef2d3 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 23:27:51 +0200 Subject: [PATCH] Issue 152: the loop is metastable, and it cleared at 23:11 It ran 22:46-23:11, five full replacements, and stopped when a pass happened to read the store during a window it was up. The record said nothing about it stops; that was wrong. Exiting by luck is the finding, not a mitigation. --- .../00-report.md | 24 +++++++++++++------ 1 file changed, 17 insertions(+), 7 deletions(-) diff --git a/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md index ccef9d9..6dcd31f 100644 --- a/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md +++ b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md @@ -11,9 +11,9 @@ amended-design: ## What was observed -For at least seventeen minutes after the last operator action, the control-node's host applied all -327 of its resources every ~6.5 minutes without pause, replacing every container on the machine each -time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail, +For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's +host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on +the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail, and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine sat near 8. Nothing was converging: each pass ended and the next began four seconds later. @@ -50,7 +50,7 @@ if err != nil { A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does not know", but "the mesh states these names do not exist", to every machine at once. -## Why it cannot recover on its own +## Why it sustains itself The loop closes through the control plane's own database: @@ -66,9 +66,19 @@ The loop closes through the control plane's own database: mesh: reporting: nats: connection closed`. 5. Back to 1. -It is stable in its instability: every pass destroys the evidence the next pass needs to decide it -has nothing to do. Nothing outside the machine has to be wrong for this to continue, and nothing -about it stops. +Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing +outside the machine has to be wrong for it to continue. + +**It is metastable, not permanent.** It ran from 22:46 to 23:11 — five full replacements of every +container on the machine — and then stopped on its own, when one pass happened to read the store +during a window it was up, composed the same roster twice running, and found nothing to do. Load fell +from 7.8 to 1.7 and the machine returned to its five-minute idle tick. + +That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not +control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or +a larger store, would not have found it. An outage that clears itself after twenty-five minutes and +five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it +is the same fault, harder to catch. ## What it is not