A seat's holders pull one ask at a time from one shared worker (hq ADR 0190, issue 186)
The worker a holder bound was a push consumer in a queue group with one ask in flight: right for one holder, and with two it would still be a queue of one — the server hands a pushed ask to whichever subscriber it picks, busy or not, and the in-flight cap is per consumer, not per holder. Now the worker is pulled: every machine holding the seat binds the same durable and fetches one ask when it has finished the last, so an idle machine is the one that takes the next, the asks in flight are bounded by the holders working, and nothing is delivered that nobody asked for — which is also what ended the race issue 186 describes. A holder's grants trade the delivery subject for MSG.NEXT on the worker; the ack grant and the heartbeat that keeps a long build alive stay. Proven against a real bus: the build round trip, a backlog taken by a machine that arrives later, and work handed back by one machine coming round again.
This commit is contained in:
@@ -126,29 +126,31 @@ func (m *natsMachine) Close() {
|
||||
}
|
||||
}
|
||||
|
||||
// Take binds to the role's worker and hands each request over, one at a time.
|
||||
// Take binds to the role's worker and pulls one request at a time, handing each over.
|
||||
//
|
||||
// **Bound, never created.** The work queue and the worker on it are the controller's to define
|
||||
// (design 25 §3), and a build machine reaches no part of the JetStream API — so a missing one is said
|
||||
// as the mesh's to answer rather than quietly created with whatever this client defaults to.
|
||||
//
|
||||
// **Pulled, one at a time, by whichever holder is free** (novox/hq ADR 0190). Every machine holding
|
||||
// the role binds this same worker; a machine asks for the next request only when it has finished
|
||||
// the last, so a slow machine never holds an ask an idle one could take, and a machine that took
|
||||
// five at once would run five container builds against one runtime and finish all of them slower
|
||||
// than the first.
|
||||
func (m *natsMachine) Take(ctx context.Context, do func(context.Context, Build)) error {
|
||||
worker, found := broker.HolderConsumerFor(m.on, "builder",
|
||||
worker, found := broker.HolderConsumerFor(m.on, "build-agent",
|
||||
broker.DeclaredSeat{Name: m.seat, Accepts: []string{"build"}})
|
||||
if !found {
|
||||
return fmt.Errorf("%s accepts no work, so there is nothing for this machine to take", m.seat)
|
||||
}
|
||||
|
||||
// One at a time, which the consumer's own ack-pending limit enforces rather than a prefetch
|
||||
// setting: a machine that took five requests at once would run five container builds against one
|
||||
// runtime and finish all of them slower than the first.
|
||||
work := make(chan *nats.Msg, 1)
|
||||
// **The consumer's own filter, not the one subject this machine cares about.** The client checks
|
||||
// what is asked for against the consumer's filter and refuses anything that is not the same —
|
||||
// "subject does not match consumer" — so subscribing `…accept.build` against a consumer filtered
|
||||
// on `…accept.>` is rejected even though it is narrower. Learned twice now, on two different
|
||||
// consumers, which is why it is written down here.
|
||||
filter := worker.Filters[0]
|
||||
sub, err := m.js.Context().ChanQueueSubscribe(filter, worker.Queue, work,
|
||||
sub, err := m.js.Context().PullSubscribe(filter, worker.Name,
|
||||
nats.Bind(worker.Stream, worker.Name), nats.ManualAck())
|
||||
if err != nil {
|
||||
return fmt.Errorf(
|
||||
@@ -159,13 +161,27 @@ func (m *natsMachine) Take(ctx context.Context, do func(context.Context, Build))
|
||||
m.sub = sub
|
||||
|
||||
for {
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
if ctx.Err() != nil {
|
||||
return nil
|
||||
case msg, ok := <-work:
|
||||
if !ok {
|
||||
return errors.New("the bus stopped delivering build work")
|
||||
}
|
||||
// One, and wait a while for it; an empty queue is a timeout, which is the normal state of a
|
||||
// machine with nothing to build, and is asked again.
|
||||
fetched, err := sub.Fetch(1, nats.Context(ctx))
|
||||
switch {
|
||||
case errors.Is(err, context.Canceled), errors.Is(err, context.DeadlineExceeded):
|
||||
return nil
|
||||
case errors.Is(err, nats.ErrTimeout):
|
||||
continue
|
||||
case err != nil:
|
||||
if sub.IsValid() {
|
||||
// A transient fault in asking — a reconnect, a slow server — is asked past rather
|
||||
// than ending the machine; one that outlasts the ack wait redelivers nothing lost.
|
||||
time.Sleep(time.Second)
|
||||
continue
|
||||
}
|
||||
return fmt.Errorf("the bus stopped delivering build work: %w", err)
|
||||
}
|
||||
for _, msg := range fetched {
|
||||
var request BuildRequest
|
||||
if err := json.Unmarshal(msg.Data, &request); err != nil {
|
||||
// Unreadable: terminated rather than retried, because the next attempt reads the same
|
||||
|
||||
Reference in New Issue
Block a user