The mesh's own roles carry a protocol, and the build branch retires

ADR 0121, first half. The `mesh-*` seats said who does a job and nothing about what
may be said to them or by them, so the mesh had roles it could not describe. They
take the same three fields a module's seat has now, and the machinery that already
derives a work queue, a holder's worker and a permission set from a declared seat
does it for these too.

The build-machine role accepts a build and emits an outcome, so `mesh.build.request`,
`mesh.control.built` and the BUILDS stream are gone. A work queue shared by several
build machines is what a seat's `accepts` already is, and keeping a second mechanism
for it was two places a permission could be wrong.

The controller's own side of a seat is a named list rather than something derived: it
is not a module and declares no `uses`, so which roles the mesh itself submits work to
has to be stated — and stating it makes that question answerable.

Two things this caught:

**The followed event subjects were hard-coded and had just gone stale.** They were
written out while the catalogue still spelled its events as the old bus's routing keys,
so converting those (issue 127) turned the pair into a controller listening to a
subject nothing publishes — the same fault as the issue, from the other side. They
derive from the emitter and the event name now, through the same function the
permission uses, so the two cannot drift apart.

**A role's queue exists before its holder**, checked against a real server, and
asserting twice changes nothing. Work queues until somebody arrives to do it, so
assigning a build machine later flushes the backlog instead of having lost it.
This commit is contained in:
2026-09-27 15:34:44 +02:00
parent 05ff6065d0
commit 0c83ecf1b5
9 changed files with 144 additions and 27 deletions
+44
View File
@@ -99,3 +99,47 @@ func TestTheControlConsumerDoesNotDeadLetterBeforeTheControllerGivesUp(t *testin
"before the controller finished deciding about it", info.Config.MaxDeliver)
}
}
// A role's work queue exists before anybody holds it, against a real server.
//
// **The queue before the holder is the point** (novox/hq ADR 0121): work queues until somebody arrives
// to do it, so assigning a build machine a week after something started asking for builds flushes the
// backlog instead of having lost it. A stream created at assignment would make "the holder is not here
// yet" mean "your requests are gone".
func TestRaisingAMeshRolesWorkQueue(t *testing.T) {
js := aLiveBus(t)
seats := []DeclaredSeat{{Name: "mesh-build-machine", Accepts: []string{"build"},
Emits: []string{"built"}}}
t.Cleanup(func() { _ = js.Context().DeleteStream("SEAT_MESH_BUILD_MACHINE") })
if err := RaiseSeats(js, seats, nil); err != nil {
t.Fatalf("a real server refused a role's work queue: %v", err)
}
info, err := js.Context().StreamInfo("SEAT_MESH_BUILD_MACHINE")
if err != nil {
t.Fatalf("the role has no work queue: %v", err)
}
if info.Config.Retention != nats.WorkQueuePolicy {
t.Errorf("the queue retains as %v: work a holder took must leave it, or the next holder does "+
"it again", info.Config.Retention)
}
if len(info.Config.Subjects) != 1 || info.Config.Subjects[0] != "mesh.seat.mesh-build-machine.accept.>" {
t.Errorf("it carries %v rather than the role's own inbound subjects", info.Config.Subjects)
}
// Nobody holds it, so there is no worker — and asserting again changes nothing, because this runs
// on every start.
if err := RaiseSeats(js, seats, nil); err != nil {
t.Fatalf("asserting a role's queue a second time failed, so a restart would: %v", err)
}
// And once somebody holds it, the worker appears on that same queue.
if err := RaiseSeats(js, seats, map[string]Holder{
"mesh-build-machine": {Node: "anchor", Module: "builder"},
}); err != nil {
t.Fatal(err)
}
if _, err := js.Context().ConsumerInfo("SEAT_MESH_BUILD_MACHINE",
"SEAT_MESH_BUILD_MACHINE_worker"); err != nil {
t.Fatalf("the holder got no worker on the role's queue: %v", err)
}
}