Issue 152: the loop is metastable, and it cleared at 23:11
It ran 22:46-23:11, five full replacements, and stopped when a pass happened to read the store during a window it was up. The record said nothing about it stops; that was wrong. Exiting by luck is the finding, not a mitigation.
This commit is contained in:
@@ -11,9 +11,9 @@ amended-design:
|
|||||||
|
|
||||||
## What was observed
|
## What was observed
|
||||||
|
|
||||||
For at least seventeen minutes after the last operator action, the control-node's host applied all
|
For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's
|
||||||
327 of its resources every ~6.5 minutes without pause, replacing every container on the machine each
|
host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on
|
||||||
time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail,
|
the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail,
|
||||||
and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine
|
and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine
|
||||||
sat near 8. Nothing was converging: each pass ended and the next began four seconds later.
|
sat near 8. Nothing was converging: each pass ended and the next began four seconds later.
|
||||||
|
|
||||||
@@ -50,7 +50,7 @@ if err != nil {
|
|||||||
A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does
|
A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does
|
||||||
not know", but "the mesh states these names do not exist", to every machine at once.
|
not know", but "the mesh states these names do not exist", to every machine at once.
|
||||||
|
|
||||||
## Why it cannot recover on its own
|
## Why it sustains itself
|
||||||
|
|
||||||
The loop closes through the control plane's own database:
|
The loop closes through the control plane's own database:
|
||||||
|
|
||||||
@@ -66,9 +66,19 @@ The loop closes through the control plane's own database:
|
|||||||
mesh: reporting: nats: connection closed`.
|
mesh: reporting: nats: connection closed`.
|
||||||
5. Back to 1.
|
5. Back to 1.
|
||||||
|
|
||||||
It is stable in its instability: every pass destroys the evidence the next pass needs to decide it
|
Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing
|
||||||
has nothing to do. Nothing outside the machine has to be wrong for this to continue, and nothing
|
outside the machine has to be wrong for it to continue.
|
||||||
about it stops.
|
|
||||||
|
**It is metastable, not permanent.** It ran from 22:46 to 23:11 — five full replacements of every
|
||||||
|
container on the machine — and then stopped on its own, when one pass happened to read the store
|
||||||
|
during a window it was up, composed the same roster twice running, and found nothing to do. Load fell
|
||||||
|
from 7.8 to 1.7 and the machine returned to its five-minute idle tick.
|
||||||
|
|
||||||
|
That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not
|
||||||
|
control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or
|
||||||
|
a larger store, would not have found it. An outage that clears itself after twenty-five minutes and
|
||||||
|
five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it
|
||||||
|
is the same fault, harder to catch.
|
||||||
|
|
||||||
## What it is not
|
## What it is not
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user