A seat's holders pull one ask at a time from one shared worker (hq ADR 0190, issue 186)

The worker a holder bound was a push consumer in a queue group with one ask in flight: right for
one holder, and with two it would still be a queue of one — the server hands a pushed ask to
whichever subscriber it picks, busy or not, and the in-flight cap is per consumer, not per holder.
Now the worker is pulled: every machine holding the seat binds the same durable and fetches one
ask when it has finished the last, so an idle machine is the one that takes the next, the asks in
flight are bounded by the holders working, and nothing is delivered that nobody asked for — which
is also what ended the race issue 186 describes. A holder's grants trade the delivery subject for
MSG.NEXT on the worker; the ack grant and the heartbeat that keeps a long build alive stay.

Proven against a real bus: the build round trip, a backlog taken by a machine that arrives later,
and work handed back by one machine coming round again.
This commit is contained in:
jochen
2026-10-02 22:34:44 +02:00
parent bde4b61b3b
commit 9f9d9b3b25
5 changed files with 80 additions and 46 deletions
+16 -15
View File
@@ -144,12 +144,18 @@ func ConsumerFor(p Principal) (Consumer, bool) {
}, true
}
// HolderConsumerFor is the worker a seat's holder gets on that seat's work queue.
// HolderConsumerFor is the worker a seat's holders share on that seat's work queue.
//
// **A queue group even though the seat guarantees one holder.** The seat is *authority* — who may
// be the telegram sender — and the queue group is *delivery*. Tie delivery to the seat and the
// day somebody allows two holders for throughput, every message is processed twice with nothing
// reporting it. Kept separate, relaxing one changes nothing about the other.
// **One worker for every holder, and each holder pulls one ask when it is idle** (novox/hq ADR
// 0190). The seat is *authority* — who may be the telegram sender — and the worker is *delivery*,
// kept separate so that relaxing one changes nothing about the other: a node-scoped seat has a
// holder per machine, and all of them take from this one consumer, so the work is shared without
// any holder knowing about the others. Pulled rather than pushed because a push consumer hands the
// next ask to whichever subscriber the server picks, busy or not, and a pulled one is asked for by
// a holder that has just become free. Which is also what ends the race issue 186 describes — asks
// delivered behind the one being worked, expiring unacknowledged and dropped after the fifth
// redelivery: nothing is delivered that nobody asked for. A long build keeps its own ask alive
// (stillWorking); the ack wait is for a holder that died.
func HolderConsumerFor(node, module string, seat DeclaredSeat) (Consumer, bool) {
if len(seat.Accepts) == 0 {
return Consumer{}, false
@@ -158,18 +164,13 @@ func HolderConsumerFor(node, module string, seat DeclaredSeat) (Consumer, bool)
Name: "SEAT_" + upperSnake(seat.Name) + "_worker",
Stream: seatStreamName(seat.Name),
Filters: []string{"mesh.seat." + seat.Name + ".accept.>"},
Queue: "holders",
AckWaitSeconds: 60,
MaxDeliver: 5,
// **One in flight.** A holder works one ask at a time, so the server hands it one at a
// time: with the default of many, every ask behind the one being worked was delivered,
// left unacknowledged for the length of the work, redelivered after the ack wait, and
// after the fifth time dropped — on 2026-10-01 twenty-six of forty-three builds asked in
// two minutes were never built, and the queue read as empty (novox/hq issue 186).
MaxAckPending: 1,
Why: fmt.Sprintf("%s on %s holds %s; it acknowledges after the work is done, so a "+
"crash mid-work redelivers rather than loses; one in flight, so a queue of asks is a "+
"queue and not a race against the ack wait", module, node, seat.Name),
// As many in flight as there are holders working, which pulling bounds by itself: a holder
// fetches one and fetches again only after it acknowledged. The server's default stands.
Why: fmt.Sprintf("%s on %s holds %s; every holder pulls one ask at a time from this worker "+
"and acknowledges after the work is done, so a crash mid-work redelivers rather than "+
"loses and an idle holder is the one that takes the next ask", module, node, seat.Name),
}, true
}