Give every test a bus of its own, at the release the mesh runs (hq ADR 0237)

The live tests reached one shared bus and assert, read and remove the mesh's own objects by
their fixed names, so packages run in parallel deleted what each other read and the suite
passed only one package at a time; a red suite read as noise. internal/testbus starts a server
per test, linked in at the nats-server release go.mod pins, and a test holds that pin to the
catalogue's bus image and to the facts snapshot's bus when there is one, so the tests never run
a bus the mesh does not. The waiter test read a timing (the most connections held at one look)
and now reads the state it means (the fewest held across the wait). make check runs the packages
in parallel under the race detector, with a timeout.
This commit is contained in:
jochen
2026-10-06 21:17:02 +02:00
parent 0090bf6af7
commit be92762969
627 changed files with 437174 additions and 166 deletions
+9 -9
View File
@@ -102,18 +102,18 @@ proxy-image:
# The whole gate. Raises a database, runs everything against it, and takes it down again --
# including when the tests fail, which is why the teardown is not conditional.
#
# **One package at a time (-p 1), and it is not about speed.** The live tests reach one bus, and on
# it they assert, read and remove the mesh's own objects -- streams and consumers with fixed names,
# because those names are the mesh's and a test cannot choose others. Two packages doing that at once
# is one deleting a consumer the other is reading through, and the failure lands in whichever test
# was reading, as "no response from stream". That reads as a bug in the code under test.
# **Packages in parallel, under the race detector, each test on a bus of its own** (internal/testbus).
# It was one package at a time against one shared bus, because the live tests assert, read and remove
# the mesh's own objects by their fixed names, and two packages at once deleted what the other read; the
# suite was red run as Go runs it and read as noise. A bus per test, of the release the mesh runs, made
# it the same in any order. The timeout bounds a hang to a failure with a stack, never a stalled gate.
check: fmt vet postgres
@go test -p 1 ./... ; status=$$? ; $(MAKE) postgres-stop ; exit $$status
@go test -race -timeout 15m ./... ; status=$$? ; $(MAKE) postgres-stop ; exit $$status
# Without a database the live tests skip rather than fail, so this is the honest subset and not
# the gate. Serialised for the same reason check is: a bus may be configured even when a store is not.
# Without a database the store's tests skip rather than fail, so this is the honest subset and not
# the gate.
test:
go test -p 1 ./...
go test -timeout 15m ./...
vet:
go vet ./...