A push consumer delivers on _DELIVER.<its name>, and a client bound to it
subscribes exactly that. No principal was granted it, and the server refused
every one the first time it bound a consumer: the control plane, each machine,
and a module would have been next. Each kind is granted its own consumers'
delivery subjects and no other's. The line announcing the raised bus printed the
URL with the credential in it; the address alone now.
Two refusals the first live connections met. A user in the MESH account was told
"JetStream not enabled for account" the first time it bound a consumer: with
accounts defined, JetStream is enabled per account, not only globally — the
account's setting, which the mesh owns, not the server's block, which it does not.
And the control plane's client used a random inbox prefix where it is granted
exactly _INBOX.<its user>.>, so the server's first answer could not reach it. The
prefix now follows from the user in the URL, for every principal that dials so.
Binding to a consumer asks the server about it and hears the answer on the
client's inbox; hearing a declaration acknowledges it. A machine's user was granted
none of that and was refused the first time one dialled a permissioned server:
"this node cannot read its declarations". Its inbox is its own prefix, which the
host now sets.
Moving the build outcome onto its role broke the one module that consumes it, and my own
agreement check passed anyway. The catalogue's subscription derived
`mesh.mod.mesh-build-machine.event.built` — a module namespace for a role's event, which no
such module owns — so it started, connected, and its graph stayed empty. The check compared
names, and the names agreed: the build machine does emit `built`. Only the subjects
disagreed, and a subscription that matches nothing is silence.
A consumed name is a module's event unless it names a role, and this package cannot tell by
looking — so whoever resolved the declaration says which, the way it already does for a seat
held or used. A module that watches a role gets the role's event subject and a consumer
filtered on it; watching grants subscribe and nothing else, because hearing what a role
announced is not taking part in it.
The check now compares the two halves that actually have to match — the subject a consumer
subscribes against the subject an emitter publishes — with a case pinning that it catches
this exact confusion. Comparing names was checking the easy half.
**And that answered the open question about catch-up: there is nothing to build.** The
mechanism exists because a queue on the old bus receives only what is published after it is
bound, so everything built before the catalogue existed was announced to nobody. A stream is
a log and a consumer is a position in it: a consumer created afterwards starts at the
beginning, so the builds are simply there. Asked of a real server, since the whole decision
rested on it — three builds published with nothing listening, then a consumer created, and
all three waiting for it.
The mesh writes its own user list, and at genesis there is no mesh yet to write it. So
the installer carries the first one — the controller's own account at a well-known
bootstrap password, exactly as the store is reached at `postgres:bootstrap` and the old
bus at `guest:guest`, and rotated with them. From the controller's first composition
onward the file is the controller's.
That left a gap I would not have found by reading: the controller's own account is
created before there is a controller to mint one, so nothing recorded a hash for it, and
its first composition would have left the writer out of the file it was writing — a bus
nothing can connect to, produced by the thing connected to it. It now records a hash of
the credential it is actually using, and only if none is recorded, so a restart cannot
put the bootstrap password back over a rotated one.
The carried list and the derived one are two statements of one fact, so a test compares
them: every subject the controller derives must be in the template, and nothing wider.
It earned itself immediately — the composer was granting both a role's whole event
branch and the one event it actually follows, which is a wider way of saying the same
thing, and the wider one wins. Only the submitting half of a role is granted now; what
comes back is named exactly.
Getting this wrong is the worst kind of silent. A controller whose carried permissions
are narrower than the ones it derives comes up, connects, and is refused on the first
thing it tries, with an authorisation error naming a subject and not the template that
forgot it — and a mesh cannot be raised twice to find out.
ADR 0121, first half. The `mesh-*` seats said who does a job and nothing about what
may be said to them or by them, so the mesh had roles it could not describe. They
take the same three fields a module's seat has now, and the machinery that already
derives a work queue, a holder's worker and a permission set from a declared seat
does it for these too.
The build-machine role accepts a build and emits an outcome, so `mesh.build.request`,
`mesh.control.built` and the BUILDS stream are gone. A work queue shared by several
build machines is what a seat's `accepts` already is, and keeping a second mechanism
for it was two places a permission could be wrong.
The controller's own side of a seat is a named list rather than something derived: it
is not a module and declares no `uses`, so which roles the mesh itself submits work to
has to be stated — and stating it makes that question answerable.
Two things this caught:
**The followed event subjects were hard-coded and had just gone stale.** They were
written out while the catalogue still spelled its events as the old bus's routing keys,
so converting those (issue 127) turned the pair into a controller listening to a
subject nothing publishes — the same fault as the issue, from the other side. They
derive from the emitter and the event name now, through the same function the
permission uses, so the two cannot drift apart.
**A role's queue exists before its holder**, checked against a real server, and
asserting twice changes nothing. Work queues until somebody arrives to do it, so
assigning a build machine later flushes the backlog instead of having lost it.
Issue 127 stood because nothing compared the two halves. Every manifest was
well-formed on its own and every derivation correct on its own, and no
cross-module subscription in the mesh matched anything — a subscription that
matches nothing is not an error, it is silence.
Two checks, because the mistake is possible at two scales.
Per manifest: an event is a local name, and `module.` is refused with the name to
write instead. A module emitting under what reads as another module's name is
refused too, pointing at the seat, where a name outlives whoever holds it.
Across the catalogue: where a consumed event's emitter is present, it must emit
that event. It cannot demand a live emitter for everything — a module lives in its
own repository and may be installed long before the one whose events it wants — so
the rule is narrower and still catches this. It found two real dangling
subscriptions the moment it ran.
Wildcards were undecided and two manifests needed them: `*` is one name and `**`
is the rest, spelled the mesh's way and derived to `>` here and `#` on the old bus.
A manifest naming either would stop being true when the wire changed, which is the
whole reason names are local.
And the field documentation taught the old form, examples included — which is why
the drift was uniform across 37 manifests rather than scattered. Nobody was
guessing; everybody followed the comment.
Found while reasoning about issue 127's replay question, in code committed earlier
today. The controller's permissions granted no inbox at all, so the answer to
every enrolment on the mesh would have been refused — "Permissions Violation for
Publish to _INBOX.enrol.anchor…" — while the controller logged that it had
enrolled the node.
**`allow_responses` does not cover it, and that is the trap.** It permits one reply
to the reply subject of a message the user received, and a message a JetStream
consumer delivers has had that field claimed for the consumer's own ack address
(design 25 §2). The address the controller actually answers is the one the request
carried in its *payload*, which the server does not recognise as a reply subject at
all. The two mechanisms look interchangeable and are not.
**My earlier verification could not have caught this.** The live enrolment tests run
against a server with no accounts and no permissions, so they exercise the subjects
and the round trip and nothing about authority. Composing the real configuration and
running a server on it is what found it.
Granted the enrolment inbox space and nothing wider: nothing but an enrolling node
ever subscribes under that prefix, each scoped to its own token's, so the controller
publishing there is the mesh answering enrolments and reaches nothing else. Confirmed
against the permissioned server both ways — the answer arrives, and a node's own
inbox is still refused.
Pinned as a rule that needs no server: whatever an enrolling node subscribes, the
controller must be able to publish to, and a node's, a module's and a person's inbox
must stay out of reach. That check is a subject-pattern match rather than a string
compare, so a grant that widened by a wildcard would not slip past it.
It also bears on 127's open question about who replays a build announcement: an
answer to a *module's* inbox would need `_INBOX.>`, which is exactly the blanket
grant design 25 §4 refuses. So the catch-up cannot become an inbox reply.
The delivery question, decided. The alternative was a manifest field enumerating
the server's ports, TLS paths and store directory so the controller could write a
whole configuration file. That is wrong: those are properties of the container the
module raises, they live in its image and its mounts, and the controller would
have to be kept in step with a Dockerfile it never sees. So the mesh writes only
what only the mesh knows — who may connect — and the module's own configuration
includes it.
`ComposeAccounts` is that file. A test says what must *not* be in it as plainly as
what must: no port, no tls block, no store_dir. Each of those in the mesh's file
is a value the controller would then own, and the module could no longer change
its own image without the mesh agreeing.
`bus-users` is where a module wants it written, and **asking is not enough to
receive it**: the file holds every user's password hash, so a module that could ask
for it could read every credential on the bus. The claim on `mesh-broker`
authorises it, checked from the manifest alone. A holder with nothing composed is
refused rather than given an empty file, for the reason a certificate is — a bus
with no user list refuses every connection in the mesh and looks like a machine
problem.
Six claims checked against a running server before any of this was committed to,
and two of them changed what got written:
**An absolute include path is resolved relative to the including file's
directory.** `include /etc/nats/accounts.conf` from /etc/nats-server/nats.conf
makes the server look for /etc/nats-server/etc/nats/accounts.conf and refuse to
start. So both files share one directory, and the module declares its own as a
file resource beside the mesh's.
**`verify: true` was refusing every connection in the mesh.** It makes the server
demand a *client* certificate, and nothing in the mesh presents one: a host pins
this server's exact certificate and authenticates with the password the mesh
minted, and so does a module's runtime. Every connection died at the TLS handshake
before any password was looked at, with an error — "client didn't provide a
certificate" — that reads as a fault in the client. Removed. TLS is still
required; verify only decides whether client certificates are checked.
The other four: a user in an included file authenticates, an unknown user is
refused so the include is the whole authority rather than an addition, a publish
outside a grant is refused, and rewriting the mesh's half alone makes a new user
appear — noticed by the module's own watcher, with no signal from outside, and
without dropping the connection the mesh already had. That last one is task 1.2's
payoff, collected.
Design 25 §6 says an enrolling node subscribes the inbox its own token derives.
It had none: `sub` was empty, so a node would publish its request and wait out
its timeout against a mesh that had answered — the handshake could not have
completed.
And there was one shared `enrolment` user, which cannot carry that inbox at all:
a permission belongs to a user, so an inbox per token means a user per token.
Named after the node, which **is** the token's id — a token is issued for a node
record, the mesh holds one live claim per record, and the node's name is the one
identifier both sides have before anything else is agreed. It is also exactly
what the other transport does, where the account is named after the node and the
secret is its password.
A nameless enrolment user is now refused rather than composed into
`_INBOX.enrol..>`: an empty subject token, and worse, one every nameless
enrolment user would share — which is one machine able to read the credentials
sealed to another.
Still to wire: something that composes one of these per live token. Nothing
composes enrolment users yet, on either bus — on the old one the account is made
imperatively through the broker's management API when a token is issued, and
here there is no management API, so issuing a token has to recompose the server's
configuration. That is the remaining half of enrolment on the new bus.
The other implementation behind the seam, so the store-window guarantee now has
both: one loop, one message at a time, the same window deciding. What differs is
where a held message lives, and that is the whole point of the move — the AMQP
side keeps an unacknowledged delivery in this process, bounded by the prefetch
and lost if the controller stops; this keeps eight bytes saying when the window
opened, and the message stays the server's.
Checked against a running server, seven claims that reasoning cannot answer: a
report is heard and leaves the work queue; one the store cannot take is naked
with a delay, stays in the stream, and is recorded when the store returns; one
about a superseded declaration is settled without being acted on; one the store
never takes is let go once the bound passes; a heartbeat is heard and nothing is
persisted; and the enrolment answer reaches the address the request carried in
its payload — the test design 25 §2 asks for, so the reason for that field
cannot quietly become folklore.
Three things the wiring forced into the open:
**The controller could not have consumed a module event.** Its permissions
granted no event subject to subscribe and no ack subject on the events stream,
so every announcement would have been redelivered for ever, refused by the list
it already had. Both narrow: each followed subject named, not `mesh.mod.*.>`.
**The controller's consumers are not derived.** It files no manifest, so its
authority cannot come from a declaration that does not exist; they sit beside the
mesh's own streams and are asserted the same way. No max-deliver on CONTROL —
the window's bound is the controller's, and a server that dead-lettered first
would discard the push the stream exists to protect.
**Channels, not callbacks.** The library would run a handler on its own
goroutine, and the window's bookkeeping is unlocked because the AMQP loop never
had two.
Design 25 §7. A person is not a module and holds no seat: nothing is
addressed to them, nothing is delivered to them, and they have no durable
consumer. What they have is permission to ask, as a list of tools or `*`
for an administrator.
Four properties the tests hold it to, each of which is a way of being
wrong that would not announce itself: a person reaches nothing but tools,
so one cannot claim a module said something; no ack subject, because
authority over a consumer that does not exist is authority nobody would
audit; no allow_responses, because a person who can answer a request is
impersonating a module on a bus where anyone may serve a tool; and two
people do not share an inbox.
Comments framed the new bus by what it replaces — a comparison in almost
every explanation, which reads as though NATS were a variant of the old
thing rather than the mesh's nervous system. Removed throughout, and
OverAMQP becomes OverCurrent: the seam's two sides are the bus the mesh
runs on today and the one being built, not two protocols.
What remains is the client library's own package name, which is its name.
Task 3.9's other half and 1.4's missing client. The derivation is pure and
unit-tested; only "does the server accept this" needs one running, behind
MESH_TEST_NATS so the ordinary suite stays offline.
A seat's work queue is created at registration, not assignment, so work
queues until a holder appears — a stream created at assignment would make
"the holder is not here yet" mean "your messages are gone". Named after the
seat, because the holder can change and the queued work must not care.
A holder's worker uses a queue group even though the seat guarantees one
holder: the seat is authority, the queue group is delivery, and tying them
together means the day somebody allows two holders every message is
processed twice with nothing reporting it.
One consumer per module carrying every filter, because its ack permission is
derived from its name.
And a real bug the live server caught: a durable name may not contain a dot,
but an ack subject is $JS.ACK.<stream>.<consumer>, so the single string that
read correctly inside the permission was rejected as a consumer name. Split
in two, beside the permission that has to match. Unfixed, the symptom would
have been every message redelivered forever with a permission list that
looks right — which is the failure design 25 §4 warns about.
Task 1.4. The foundation set only — a seat's streams come at registration
and a module's consumers at assignment, neither of which has happened at
genesis (ADR 0118).
Asserted rather than created: a stream that was deleted, or a mesh raised
from a backup, must converge rather than run without the guarantee its
messages assume.
Two things the definitions have to get right, both tested:
- CONTROL names its subjects instead of taking mesh.control.>, because
heartbeats live under that prefix and a stream of them competes for
retention with the messages that matter
- EVENTS filters on the event token, which is why that token exists; a
filter over a module's whole namespace would persist every tool call
Overlapping filters are refused where the set is written: NATS accepts two
streams matching one subject and stores the message twice under two
retentions, which nothing reports.
Adds nats.go as a dependency; it pulled golang.org/x/* forward. Full suite
green.
Task 1.3 of novox/hq ADR 0116. On AMQP an account was an HTTP call; on NATS
it is text the controller composes and the server reloads (ADR 0106). Pure,
so the mesh's whole authority model is testable as strings.
NATS closes a gap management.go recorded rather than hid: LavinMQ has no
topic permissions, so an emitter was granted the events exchange whole and
ADR 0042's origin reservation was "stamped by the sdk, not enforced here".
Per-subject permissions make it the server's refusal.
Two things found by composing a real file rather than reading the design:
- a scoped inbox leaves a responder unable to reply, because the answer goes
to the caller's inbox. allow_responses is the answer — one reply to the
subject of a message actually received — and only principals that serve
are granted it. Recorded in design 25 §4.
- composition must be deterministic: the module's entrypoint reloads on the
file's digest, so an order-dependent composer would reload the whole bus
on every controller restart. Covered by a test.
The golden fixture is the exact text `nats-server -t` accepts, so the syntax
is the server's rather than one we invented.