Issue 230: a report is also lost when the apply restarts the bus
This commit is contained in:
+9
-2
@@ -37,6 +37,12 @@ Nothing logged, alerted or counted the wait. It was found because a person asked
|
||||
state, and the cause was found by reading a machine's own journal. A push to each of the three machines
|
||||
released it: each new host applied and reported, and the plan moved on.
|
||||
|
||||
The same day, a second way to lose a report showed up. Assigning modules with tools to a workstation
|
||||
changed the bus's user list, which the control machine's declaration carries. Applying it replaced the
|
||||
bus's container, which cut every machine off for about fifteen seconds. The control machine itself
|
||||
then logged `applied, and could not tell the mesh: reporting: nats: connection closed`. The report was
|
||||
lost because the bus restarted under the apply that restarted it.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
**Every genuine host upgrade loses one report.**
|
||||
@@ -60,8 +66,9 @@ answer takes, something is wrong, and the mesh must say so itself.
|
||||
|
||||
## What a fix has to settle
|
||||
|
||||
1. **The host reports before it stands aside.** The report should be published and acknowledged
|
||||
before the old host exits. Failing that, the new host should report the declaration it took over,
|
||||
1. **A report survives whatever its own apply restarts.** The host publishes its report and has it
|
||||
acknowledged before it stands aside. It retries a report the bus dropped once the link is back.
|
||||
Failing that, the new host should report the declaration it took over,
|
||||
naming the apply its predecessor finished. A lost report must be impossible, not merely unlikely.
|
||||
2. **A plan's wait has an age and a bound.**
|
||||
- "for 0s" must be the real time since the wait began.
|
||||
|
||||
Reference in New Issue
Block a user