Commit Graph
3 Commits
Author SHA1 Message Date
jschoubben 0c83ecf1b5 The mesh's own roles carry a protocol, and the build branch retires
ADR 0121, first half. The `mesh-*` seats said who does a job and nothing about what
may be said to them or by them, so the mesh had roles it could not describe. They
take the same three fields a module's seat has now, and the machinery that already
derives a work queue, a holder's worker and a permission set from a declared seat
does it for these too.

The build-machine role accepts a build and emits an outcome, so `mesh.build.request`,
`mesh.control.built` and the BUILDS stream are gone. A work queue shared by several
build machines is what a seat's `accepts` already is, and keeping a second mechanism
for it was two places a permission could be wrong.

The controller's own side of a seat is a named list rather than something derived: it
is not a module and declares no `uses`, so which roles the mesh itself submits work to
has to be stated — and stating it makes that question answerable.

Two things this caught:

**The followed event subjects were hard-coded and had just gone stale.** They were
written out while the catalogue still spelled its events as the old bus's routing keys,
so converting those (issue 127) turned the pair into a controller listening to a
subject nothing publishes — the same fault as the issue, from the other side. They
derive from the emitter and the event name now, through the same function the
permission uses, so the two cannot drift apart.

**A role's queue exists before its holder**, checked against a real server, and
asserting twice changes nothing. Work queues until somebody arrives to do it, so
assigning a build machine later flushes the backlog instead of having lost it.
2026-09-27 15:34:44 +02:00
jschoubben eb72ec36ba 1.7 finished: minting, the file delivered, and a test flake I caused
**First, a correction: the previous commit went in on a false check.** Its message
says the suite passed; it did not. The check piped `go test` through a filter that
swallowed the failures and then printed "green" regardless. Two tests were failing
when 4de10e3 landed.

What was failing was my own doing. Purging the streams instead of deleting them
(4de10e3) left the *consumers* behind, because deleting a stream takes its
consumers with it and purging does not. A durable push consumer surviving between
tests keeps pushing to a delivery subject the previous test's subscription has gone
from: the messages count as delivered, go nowhere, and the next test waits out its
timeout for an announcement the server believes it already sent. Consumers are now
removed with the purge. Five consecutive clean runs.

`-p 1` stays, because two packages asserting and deleting the same fixed-name
objects on one bus is a real race — but its comment said the cause I had guessed
and not the one I found, so it now says the right thing.

**And delivery was not finished when I said it was.** Nothing filled
`Rendering.BusUsers`, so the composed file would never have reached a node.
`composeBusUsers` closes it: composed per push for the machine holding
`mesh-broker`, never kept, because the list is a function of the mesh's records and
a stored copy could disagree with them while both looked consistent. A user with no
credential is left out and named rather than written as a user without a password —
an ordinary situation with an obvious remedy — but a file with no users at all is
refused, because that bus would refuse every connection in the mesh.

**Minting, on both halves.** A node at enrolment and a module at `module issue`.
Three things differ from a management call and each is the point of the move: the
credential is minted into the mesh's records and becomes usable at the next
composition, so no server need be reachable; the password travels beside the address
rather than inside it, because a credential embedded in a URL leaks into every log
line that prints a connection; and a module's durable consumer is derived from what
it declared rather than named, so it cannot ask for delivery of something it did not
say it consumes.

A node reconnecting may be refused until that composition reaches the machine
running the bus. That is what the host's reconnect backoff is for and it is
survivable by design; waiting for the push would hold an enrolment open for as long
as a declaration takes to apply.

Tested that the switch is a switch: a node enrolling on one bus comes away with a
credential for that bus and none for the other, because one that held both could be
half-moved and nothing would say which half.
2026-09-27 03:19:41 +02:00
jschoubben 4de10e32e3 The bus's objects are raised on every start, and one switch says which bus
Two of 1.7's three remaining pieces.

**Raised on every start, not created once at genesis.** A stream somebody deleted,
a mesh raised from a restored backup, or a bus whose data directory was replaced
all have records and no objects — and a node whose consumer is missing hears
nothing while everything else about it looks correct.

The order is not a preference: a consumer on a stream that does not exist is
refused *naming the stream*, so somebody reading that refusal goes looking for a
deletion instead of a reversed pair of lines. Pinned by a test, along with the one
thing about seats that reads like an omission and is not — a seat's work queue is
asserted whether or not anybody holds it, because work queues until a holder
appears, so installing the module a week later flushes the backlog instead of
having lost it.

Against a real server: every object accepted, asserting twice changes nothing (a
start that failed the second time is a controller that cannot restart), a machine
joining an already-raised bus is accepted, each node's consumer is bound to its own
declaration subject and no other's, and CONTROL does not dead-letter — because the
store window's bound belongs to the controller and a server that gave up first
would discard the push the stream exists to protect.

**Which bus this mesh is on is one fact, read in one place.** Every seam the change
went behind ships both implementations; this is what the rollout flips. Being told
about both is refused at start rather than warned about: a mesh half on each is one
where a declaration goes out on one bus and the report comes back on the other, and
every component logs success while it happens — ADR 0074's failure arriving through
configuration instead of through code. The refusal names both variables and says
which to unset, because whoever reads it has to choose and the wrong choice is a
rollout half done.
2026-09-27 02:59:05 +02:00