Commit Graph
5 Commits
Author SHA1 Message Date
jschoubben 0c83ecf1b5 The mesh's own roles carry a protocol, and the build branch retires
ADR 0121, first half. The `mesh-*` seats said who does a job and nothing about what
may be said to them or by them, so the mesh had roles it could not describe. They
take the same three fields a module's seat has now, and the machinery that already
derives a work queue, a holder's worker and a permission set from a declared seat
does it for these too.

The build-machine role accepts a build and emits an outcome, so `mesh.build.request`,
`mesh.control.built` and the BUILDS stream are gone. A work queue shared by several
build machines is what a seat's `accepts` already is, and keeping a second mechanism
for it was two places a permission could be wrong.

The controller's own side of a seat is a named list rather than something derived: it
is not a module and declares no `uses`, so which roles the mesh itself submits work to
has to be stated — and stating it makes that question answerable.

Two things this caught:

**The followed event subjects were hard-coded and had just gone stale.** They were
written out while the catalogue still spelled its events as the old bus's routing keys,
so converting those (issue 127) turned the pair into a controller listening to a
subject nothing publishes — the same fault as the issue, from the other side. They
derive from the emitter and the event name now, through the same function the
permission uses, so the two cannot drift apart.

**A role's queue exists before its holder**, checked against a real server, and
asserting twice changes nothing. Work queues until somebody arrives to do it, so
assigning a build machine later flushes the backlog instead of having lost it.
2026-09-27 15:34:44 +02:00
jschoubben 05ff6065d0 Event names are checked now, per manifest and across the catalogue
Issue 127 stood because nothing compared the two halves. Every manifest was
well-formed on its own and every derivation correct on its own, and no
cross-module subscription in the mesh matched anything — a subscription that
matches nothing is not an error, it is silence.

Two checks, because the mistake is possible at two scales.

Per manifest: an event is a local name, and `module.` is refused with the name to
write instead. A module emitting under what reads as another module's name is
refused too, pointing at the seat, where a name outlives whoever holds it.

Across the catalogue: where a consumed event's emitter is present, it must emit
that event. It cannot demand a live emitter for everything — a module lives in its
own repository and may be installed long before the one whose events it wants — so
the rule is narrower and still catches this. It found two real dangling
subscriptions the moment it ran.

Wildcards were undecided and two manifests needed them: `*` is one name and `**`
is the rest, spelled the mesh's way and derived to `>` here and `#` on the old bus.
A manifest naming either would stop being true when the wire changed, which is the
whole reason names are local.

And the field documentation taught the old form, examples included — which is why
the drift was uniform across 37 manifests rather than scattered. Nobody was
guessing; everybody followed the comment.
2026-09-27 14:43:16 +02:00
jschoubben e65b3950cc A node's declaration consumer, which only the mesh can make
A host's account reaches no part of the JetStream API — correctly, because the
controller is the only writer of consumer definitions — so the object a node
reads its declarations through has to be waiting before the host binds to it,
and nothing created one. Named after the node, because the node's own ack grant
is `$JS.ACK.NODES.<node>.>` and a consumer named anything else is one the host
cannot acknowledge a delivery from.

No max-deliver, and a five-minute ack wait: a declaration is settled only after
the node has applied it and reported, which is minutes on a machine pulling
images, and the stream holds exactly one message per node — so there is nothing
to dead-letter, only one message to redeliver for as long as that node is away.

Asserted on start as well as created at enrolment, for the reason the streams
are: a mesh raised from a restored backup has node records and no consumers, and
a node whose consumer is missing hears nothing while everything else about it
looks correct.

The test sets the consumer against the grant the node actually gets, because
each of the three ways of getting it wrong is silent: a wrong name cannot ack, a
wrong filter reads another node's declarations, and a pull consumer is one a host
has no authority to bind.
2026-09-27 01:24:46 +02:00
jschoubben 88bef39952 The consume side on NATS, and the window held by the server
The other implementation behind the seam, so the store-window guarantee now has
both: one loop, one message at a time, the same window deciding. What differs is
where a held message lives, and that is the whole point of the move — the AMQP
side keeps an unacknowledged delivery in this process, bounded by the prefetch
and lost if the controller stops; this keeps eight bytes saying when the window
opened, and the message stays the server's.

Checked against a running server, seven claims that reasoning cannot answer: a
report is heard and leaves the work queue; one the store cannot take is naked
with a delay, stays in the stream, and is recorded when the store returns; one
about a superseded declaration is settled without being acted on; one the store
never takes is let go once the bound passes; a heartbeat is heard and nothing is
persisted; and the enrolment answer reaches the address the request carried in
its payload — the test design 25 §2 asks for, so the reason for that field
cannot quietly become folklore.

Three things the wiring forced into the open:

**The controller could not have consumed a module event.** Its permissions
granted no event subject to subscribe and no ack subject on the events stream,
so every announcement would have been redelivered for ever, refused by the list
it already had. Both narrow: each followed subject named, not `mesh.mod.*.>`.

**The controller's consumers are not derived.** It files no manifest, so its
authority cannot come from a declaration that does not exist; they sit beside the
mesh's own streams and are asserted the same way. No max-deliver on CONTROL —
the window's bound is the controller's, and a server that dead-lettered first
would discard the push the stream exists to protect.

**Channels, not callbacks.** The library would run a handler on its own
goroutine, and the window's bookkeeping is unlocked because the AMQP loop never
had two.
2026-09-27 00:53:33 +02:00
jschoubben aa74bd86ca Derive a seat's stream and a module's consumer, and wire JetStream
Task 3.9's other half and 1.4's missing client. The derivation is pure and
unit-tested; only "does the server accept this" needs one running, behind
MESH_TEST_NATS so the ordinary suite stays offline.

A seat's work queue is created at registration, not assignment, so work
queues until a holder appears — a stream created at assignment would make
"the holder is not here yet" mean "your messages are gone". Named after the
seat, because the holder can change and the queued work must not care.

A holder's worker uses a queue group even though the seat guarantees one
holder: the seat is authority, the queue group is delivery, and tying them
together means the day somebody allows two holders every message is
processed twice with nothing reporting it.

One consumer per module carrying every filter, because its ack permission is
derived from its name.

And a real bug the live server caught: a durable name may not contain a dot,
but an ack subject is $JS.ACK.<stream>.<consumer>, so the single string that
read correctly inside the permission was rejected as a consumer name. Split
in two, beside the permission that has to match. Unfixed, the symptom would
have been every message redelivered forever with a permission list that
looks right — which is the failure design 25 §4 warns about.
2026-09-26 22:28:42 +02:00