The build machine takes one ask at a time, and says so while it builds
With the worker consumer's default of many deliveries in flight, every ask behind the one being built was delivered at once, left unacknowledged for the length of the build, redelivered after the ack wait and dropped after the fifth time: on 2026-10-01 twenty-six of forty-three builds asked in two minutes were never built and the queue read as empty (hq issue 186). The holder's worker now has one in flight, and a running build tells the bus it is still working, as the controller's long handlers do, so a build longer than the ack wait is neither redelivered nor counted out.
This commit is contained in:
@@ -161,8 +161,15 @@ func HolderConsumerFor(node, module string, seat DeclaredSeat) (Consumer, bool)
|
||||
Queue: "holders",
|
||||
AckWaitSeconds: 60,
|
||||
MaxDeliver: 5,
|
||||
// **One in flight.** A holder works one ask at a time, so the server hands it one at a
|
||||
// time: with the default of many, every ask behind the one being worked was delivered,
|
||||
// left unacknowledged for the length of the work, redelivered after the ack wait, and
|
||||
// after the fifth time dropped — on 2026-10-01 twenty-six of forty-three builds asked in
|
||||
// two minutes were never built, and the queue read as empty (novox/hq issue 186).
|
||||
MaxAckPending: 1,
|
||||
Why: fmt.Sprintf("%s on %s holds %s; it acknowledges after the work is done, so a "+
|
||||
"crash mid-work redelivers rather than loses", module, node, seat.Name),
|
||||
"crash mid-work redelivers rather than loses; one in flight, so a queue of asks is a "+
|
||||
"queue and not a race against the ack wait", module, node, seat.Name),
|
||||
}, true
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user