The outbound half went behind `Bus` and the transport stopped reaching its callers; this is the other half, and the larger one. Every handler took `amqp.Delivery`, so the serving loop could not move to another bus without moving enrolment, reports, builds, upgrades and catch-up with it in one breath. `Control` states one message in the mesh's words — took it, dropped it, or held it for the store — and `Inbound` is where messages come from. The AMQP implementation is today's loop moved rather than changed: same queues, same prefetch, same holding, because the mesh is running on it and a bus nothing speaks yet is no reason to alter the one every node is on. The window (window.go) is now what decides, instead of the conditions that were inlined in the loop. Two things that surfaced in the wiring: **Supersession is asked before the store, not after.** A report about a declaration the mesh has moved past would otherwise wait out a restarting store to be written and then overwrite what the node is doing now. **Half of a report is not about a declaration, and that half is never stale.** What the machine *is* — the tunnel it took over, the ports its own bundle holds, what an adopted node found, a node moving its overlay key — reaches the mesh on a report and nowhere else. A rekey set aside as stale is a node whose overlay key never moves, and no retry is coming, because the node said it once. So staleness is asked only of a report that is purely an apply's account. The one thing holding-in-memory can do that holding-in-the-server cannot is named rather than hidden: `About` sets aside a held message when a newer one about the same thing arrives, and the bus being built ignores it because the digest answers the same question.
121 lines
5.5 KiB
Go
121 lines
5.5 KiB
Go
package link
|
|
|
|
import (
|
|
"errors"
|
|
"time"
|
|
|
|
"github.com/novox/mesh-controller/internal/inventory"
|
|
)
|
|
|
|
// The store window, as a decision rather than a mechanism.
|
|
//
|
|
// The guarantee (novox/hq ADR 0083): a push the controller cannot record because its store is
|
|
// restarting is **held and retried**, not dropped and not falsely acknowledged. On the bus the
|
|
// mesh runs on today that is done by keeping the delivery unacknowledged in memory and settling
|
|
// it later. On the bus being built it is a `nak` with a delay: the server holds it and redelivers,
|
|
// so the controller keeps no list of parked messages and a controller that restarts mid-window
|
|
// loses nothing it was holding.
|
|
//
|
|
// **The decision is the same either way, and the mechanism is not the interesting part.** What is
|
|
// interesting is that moving the holding into the server introduces a problem the in-memory
|
|
// version did not have, and the answer was already in the message.
|
|
|
|
// Verdict is what to do with one control message.
|
|
type Verdict int
|
|
|
|
const (
|
|
// Take it: apply, then acknowledge.
|
|
Take Verdict = iota
|
|
// Hold it: the store cannot record this yet. Nak with a delay and let the server redeliver.
|
|
Hold
|
|
// Stale: this is about a declaration the node has already moved past, and applying it would
|
|
// undo what came after. Acknowledge without acting — redelivering forever is worse.
|
|
Stale
|
|
// GiveUp: the store has not come back in time. Settle it and say so, loudly.
|
|
GiveUp
|
|
)
|
|
|
|
// StoreWindow decides. Pure, so the guarantee is testable without a bus, a store or a clock.
|
|
type StoreWindow struct {
|
|
// GiveUpAfter is how long one message may be held before it is let go with a line saying so.
|
|
GiveUpAfter time.Duration
|
|
}
|
|
|
|
// Decide answers for one delivery.
|
|
//
|
|
// - err is what the store said, or nil.
|
|
// - declaredIn is the digest of the declaration this message is about, empty when it is not
|
|
// about one (an enrolment, a build result).
|
|
// - outstanding is the digest the mesh last sent that node, empty when it has sent none.
|
|
// - heldFor is how long this message has already been held; zero on first delivery.
|
|
func (w StoreWindow) Decide(err error, declaredIn, outstanding string, heldFor time.Duration) Verdict {
|
|
// **Staleness is checked before the store, not after.** A redelivery that lost its race is
|
|
// not worth waiting on a store for, and asking the store first would mean a message about a
|
|
// superseded declaration holding a slot in the window that a current one needs.
|
|
if Superseded(declaredIn, outstanding) {
|
|
return Stale
|
|
}
|
|
if err == nil {
|
|
return Take
|
|
}
|
|
if !errors.Is(err, ErrTryAgain) && !inventory.Unreachable(err) {
|
|
// Not the store being away: a refusal is an answer, and holding it would turn a message
|
|
// the mesh understood into one it retries forever.
|
|
return Take
|
|
}
|
|
if heldFor >= w.GiveUpAfter {
|
|
return GiveUp
|
|
}
|
|
return Hold
|
|
}
|
|
|
|
// Superseded says a message is about a declaration the mesh has already moved past.
|
|
//
|
|
// Stated on its own because it is asked in two places for one reason: here, so the window's whole
|
|
// decision is in one pure function, and by the serving loop *before* it asks the store, because
|
|
// that is the point — a message about the past must not wait on a store, or it holds a slot in the
|
|
// window that a current message needs.
|
|
//
|
|
// Unanswerable is not stale. A message that names no declaration, and a node the mesh has never
|
|
// sent one, both give nothing to compare: the mesh acts on the message rather than guessing, which
|
|
// is also what keeps a host built before reports carried the digest from going silent.
|
|
func Superseded(declaredIn, outstanding string) bool {
|
|
return declaredIn != "" && outstanding != "" && declaredIn != outstanding
|
|
}
|
|
|
|
// RedeliverAfter is how long the server should hold a naked message before trying again.
|
|
//
|
|
// Backed off, and bounded. A store restarting is back in seconds; a store that is gone is not
|
|
// helped by being asked every second, and the delay is what keeps a window of held messages from
|
|
// becoming a spin.
|
|
func RedeliverAfter(heldFor time.Duration) time.Duration {
|
|
switch {
|
|
case heldFor < 5*time.Second:
|
|
return time.Second
|
|
case heldFor < 30*time.Second:
|
|
return 5 * time.Second
|
|
default:
|
|
return 15 * time.Second
|
|
}
|
|
}
|
|
|
|
// The problem holding-in-the-server introduces, and why the answer was already in the message.
|
|
//
|
|
// Holding a delivery in memory let the controller do something a server cannot: when a newer
|
|
// report for the same node arrived, it dropped the older one, "because acting on it after the
|
|
// newer would undo the newer". A `nak`ed message is the server's, and the server will redeliver
|
|
// it whatever else has happened in the meantime — so the older report comes back *after* the
|
|
// newer was applied, and applying it would undo exactly what that comment describes.
|
|
//
|
|
// **A report already says which declaration it is about.** `Declared` is the digest of the exact
|
|
// bytes the mesh sent, and it exists because an earlier attempt to order reports by time lost the
|
|
// race it invited — an apply that started under the previous declaration finishes after the next
|
|
// is sent, and the report reads as newer than the send. Clocks cannot answer *which*; the digest
|
|
// is the answer itself.
|
|
//
|
|
// So supersession stops being a thing the controller remembers and becomes a thing it checks: a
|
|
// report whose digest is not the one outstanding for that node is stale, and is acknowledged
|
|
// without being acted on. Which is the same shape as a node refusing a superseded declaration by
|
|
// sequence (novox/hq issue 107) — ordering settled by what the message says, not by when it
|
|
// happened to arrive.
|