Commit Graph
9 Commits
Author SHA1 Message Date
jschoubben 8e2824201a Genesis can raise a mesh on the new bus, and the carried user list is checked against the composer
The mesh writes its own user list, and at genesis there is no mesh yet to write it. So
the installer carries the first one — the controller's own account at a well-known
bootstrap password, exactly as the store is reached at `postgres:bootstrap` and the old
bus at `guest:guest`, and rotated with them. From the controller's first composition
onward the file is the controller's.

That left a gap I would not have found by reading: the controller's own account is
created before there is a controller to mint one, so nothing recorded a hash for it, and
its first composition would have left the writer out of the file it was writing — a bus
nothing can connect to, produced by the thing connected to it. It now records a hash of
the credential it is actually using, and only if none is recorded, so a restart cannot
put the bootstrap password back over a rotated one.

The carried list and the derived one are two statements of one fact, so a test compares
them: every subject the controller derives must be in the template, and nothing wider.
It earned itself immediately — the composer was granting both a role's whole event
branch and the one event it actually follows, which is a wider way of saying the same
thing, and the wider one wins. Only the submitting half of a role is granted now; what
comes back is named exactly.

Getting this wrong is the worst kind of silent. A controller whose carried permissions
are narrower than the ones it derives comes up, connects, and is refused on the first
thing it tries, with an authorisation error naming a subject and not the template that
forgot it — and a mesh cannot be raised twice to find out.
2026-09-27 16:39:19 +02:00
jschoubben e4e960ec1c A build is work submitted to a role, on both buses
ADR 0121 carried through to working code. `Builders` is the asking side and
`BuildMachine` the taking side, each with an implementation per bus, and the builder
binary and the `build` command now go through them.

On the bus being built, one publish does what two did. The old bus answered the asker
through a reply queue and announced to an events exchange, because two audiences meant
two topologies. Here the outcome is the role's own event: the asker matches it by the
id its request carried, the controller records it, the catalogue places it in the graph.
So a build machine publishes once, needs a reply queue for nothing, and needs a grant
over nobody's inbox — which is what ruled out the alternatives.

The outcome carries the module name now. Only the manifest says what was built, and on
the old bus the separate announcement carried it; with one message for three readers it
belongs in the result. A failed build names none, because it produced no module version
and the catalogue would otherwise place something that was never made.

Checked against a real server: the whole round trip; a third party on the role's event
hearing the same outcome the asker did, which is the claim the decision rests on; work
leaving the queue once settled, so no second machine repeats it; work submitted with no
machine holding the role waiting instead of failing, and being done when one arrives;
and work a machine handed back coming round again.

One thing I got wrong twice now and have written down where it bit: binding to a
consumer must name that consumer's own filter subject, not the narrower subject the
caller cares about. The client compares the two and refuses anything that is not equal,
with "subject does not match consumer".
2026-09-27 16:01:53 +02:00
jschoubben 0c83ecf1b5 The mesh's own roles carry a protocol, and the build branch retires
ADR 0121, first half. The `mesh-*` seats said who does a job and nothing about what
may be said to them or by them, so the mesh had roles it could not describe. They
take the same three fields a module's seat has now, and the machinery that already
derives a work queue, a holder's worker and a permission set from a declared seat
does it for these too.

The build-machine role accepts a build and emits an outcome, so `mesh.build.request`,
`mesh.control.built` and the BUILDS stream are gone. A work queue shared by several
build machines is what a seat's `accepts` already is, and keeping a second mechanism
for it was two places a permission could be wrong.

The controller's own side of a seat is a named list rather than something derived: it
is not a module and declares no `uses`, so which roles the mesh itself submits work to
has to be stated — and stating it makes that question answerable.

Two things this caught:

**The followed event subjects were hard-coded and had just gone stale.** They were
written out while the catalogue still spelled its events as the old bus's routing keys,
so converting those (issue 127) turned the pair into a controller listening to a
subject nothing publishes — the same fault as the issue, from the other side. They
derive from the emitter and the event name now, through the same function the
permission uses, so the two cannot drift apart.

**A role's queue exists before its holder**, checked against a real server, and
asserting twice changes nothing. Work queues until somebody arrives to do it, so
assigning a build machine later flushes the backlog instead of having lost it.
2026-09-27 15:34:44 +02:00
jschoubben d65c37caad The controller could not answer an enrolment, and a probe on an open server said it could
Found while reasoning about issue 127's replay question, in code committed earlier
today. The controller's permissions granted no inbox at all, so the answer to
every enrolment on the mesh would have been refused — "Permissions Violation for
Publish to _INBOX.enrol.anchor…" — while the controller logged that it had
enrolled the node.

**`allow_responses` does not cover it, and that is the trap.** It permits one reply
to the reply subject of a message the user received, and a message a JetStream
consumer delivers has had that field claimed for the consumer's own ack address
(design 25 §2). The address the controller actually answers is the one the request
carried in its *payload*, which the server does not recognise as a reply subject at
all. The two mechanisms look interchangeable and are not.

**My earlier verification could not have caught this.** The live enrolment tests run
against a server with no accounts and no permissions, so they exercise the subjects
and the round trip and nothing about authority. Composing the real configuration and
running a server on it is what found it.

Granted the enrolment inbox space and nothing wider: nothing but an enrolling node
ever subscribes under that prefix, each scoped to its own token's, so the controller
publishing there is the mesh answering enrolments and reaches nothing else. Confirmed
against the permissioned server both ways — the answer arrives, and a node's own
inbox is still refused.

Pinned as a rule that needs no server: whatever an enrolling node subscribes, the
controller must be able to publish to, and a node's, a module's and a person's inbox
must stay out of reach. That check is a subject-pattern match rather than a string
compare, so a grant that widened by a wildcard would not slip past it.

It also bears on 127's open question about who replays a build announcement: an
answer to a *module's* inbox would need `_INBOX.>`, which is exactly the blanket
grant design 25 §4 refuses. So the catch-up cannot become an inbox reply.
2026-09-27 14:13:52 +02:00
jschoubben f8ab9f2dcf The mesh composes the accounts; the module owns its server
The delivery question, decided. The alternative was a manifest field enumerating
the server's ports, TLS paths and store directory so the controller could write a
whole configuration file. That is wrong: those are properties of the container the
module raises, they live in its image and its mounts, and the controller would
have to be kept in step with a Dockerfile it never sees. So the mesh writes only
what only the mesh knows — who may connect — and the module's own configuration
includes it.

`ComposeAccounts` is that file. A test says what must *not* be in it as plainly as
what must: no port, no tls block, no store_dir. Each of those in the mesh's file
is a value the controller would then own, and the module could no longer change
its own image without the mesh agreeing.

`bus-users` is where a module wants it written, and **asking is not enough to
receive it**: the file holds every user's password hash, so a module that could ask
for it could read every credential on the bus. The claim on `mesh-broker`
authorises it, checked from the manifest alone. A holder with nothing composed is
refused rather than given an empty file, for the reason a certificate is — a bus
with no user list refuses every connection in the mesh and looks like a machine
problem.

Six claims checked against a running server before any of this was committed to,
and two of them changed what got written:

**An absolute include path is resolved relative to the including file's
directory.** `include /etc/nats/accounts.conf` from /etc/nats-server/nats.conf
makes the server look for /etc/nats-server/etc/nats/accounts.conf and refuse to
start. So both files share one directory, and the module declares its own as a
file resource beside the mesh's.

**`verify: true` was refusing every connection in the mesh.** It makes the server
demand a *client* certificate, and nothing in the mesh presents one: a host pins
this server's exact certificate and authenticates with the password the mesh
minted, and so does a module's runtime. Every connection died at the TLS handshake
before any password was looked at, with an error — "client didn't provide a
certificate" — that reads as a fault in the client. Removed. TLS is still
required; verify only decides whether client certificates are checked.

The other four: a user in an included file authenticates, an unknown user is
refused so the include is the whole authority rather than an addition, a publish
outside a grant is refused, and rewriting the mesh's half alone makes a new user
appear — noticed by the module's own watcher, with no signal from outside, and
without dropping the connection the mesh already had. That last one is task 1.2's
payoff, collected.
2026-09-27 02:50:23 +02:00
jschoubben 7180a273a2 The enrolment user is per token, and it has an inbox
Design 25 §6 says an enrolling node subscribes the inbox its own token derives.
It had none: `sub` was empty, so a node would publish its request and wait out
its timeout against a mesh that had answered — the handshake could not have
completed.

And there was one shared `enrolment` user, which cannot carry that inbox at all:
a permission belongs to a user, so an inbox per token means a user per token.
Named after the node, which **is** the token's id — a token is issued for a node
record, the mesh holds one live claim per record, and the node's name is the one
identifier both sides have before anything else is agreed. It is also exactly
what the other transport does, where the account is named after the node and the
secret is its password.

A nameless enrolment user is now refused rather than composed into
`_INBOX.enrol..>`: an empty subject token, and worse, one every nameless
enrolment user would share — which is one machine able to read the credentials
sealed to another.

Still to wire: something that composes one of these per live token. Nothing
composes enrolment users yet, on either bus — on the old one the account is made
imperatively through the broker's management API when a token is issued, and
here there is no management API, so issuing a token has to recompose the server's
configuration. That is the remaining half of enrolment on the new bus.
2026-09-27 01:31:34 +02:00
jschoubben 88bef39952 The consume side on NATS, and the window held by the server
The other implementation behind the seam, so the store-window guarantee now has
both: one loop, one message at a time, the same window deciding. What differs is
where a held message lives, and that is the whole point of the move — the AMQP
side keeps an unacknowledged delivery in this process, bounded by the prefetch
and lost if the controller stops; this keeps eight bytes saying when the window
opened, and the message stays the server's.

Checked against a running server, seven claims that reasoning cannot answer: a
report is heard and leaves the work queue; one the store cannot take is naked
with a delay, stays in the stream, and is recorded when the store returns; one
about a superseded declaration is settled without being acted on; one the store
never takes is let go once the bound passes; a heartbeat is heard and nothing is
persisted; and the enrolment answer reaches the address the request carried in
its payload — the test design 25 §2 asks for, so the reason for that field
cannot quietly become folklore.

Three things the wiring forced into the open:

**The controller could not have consumed a module event.** Its permissions
granted no event subject to subscribe and no ack subject on the events stream,
so every announcement would have been redelivered for ever, refused by the list
it already had. Both narrow: each followed subject named, not `mesh.mod.*.>`.

**The controller's consumers are not derived.** It files no manifest, so its
authority cannot come from a declaration that does not exist; they sit beside the
mesh's own streams and are asserted the same way. No max-deliver on CONTROL —
the window's bound is the controller's, and a server that dead-lettered first
would discard the push the stream exists to protect.

**Channels, not callbacks.** The library would run a handler on its own
goroutine, and the window's bookkeeping is unlocked because the AMQP loop never
had two.
2026-09-27 00:53:33 +02:00
jschoubben 1f36787d75 The mesh's four streams, asserted on every start
Task 1.4. The foundation set only — a seat's streams come at registration
and a module's consumers at assignment, neither of which has happened at
genesis (ADR 0118).

Asserted rather than created: a stream that was deleted, or a mesh raised
from a backup, must converge rather than run without the guarantee its
messages assume.

Two things the definitions have to get right, both tested:
- CONTROL names its subjects instead of taking mesh.control.>, because
  heartbeats live under that prefix and a stream of them competes for
  retention with the messages that matter
- EVENTS filters on the event token, which is why that token exists; a
  filter over a module's whole namespace would persist every tool call

Overlapping filters are refused where the set is written: NATS accepts two
streams matching one subject and stores the message twice under two
retentions, which nothing reports.

Adds nats.go as a dependency; it pulled golang.org/x/* forward. Full suite
green.
2026-09-26 21:02:18 +02:00
jschoubben c753f9d5c0 Compose the bus's accounts instead of calling a management API
Task 1.3 of novox/hq ADR 0116. On AMQP an account was an HTTP call; on NATS
it is text the controller composes and the server reloads (ADR 0106). Pure,
so the mesh's whole authority model is testable as strings.

NATS closes a gap management.go recorded rather than hid: LavinMQ has no
topic permissions, so an emitter was granted the events exchange whole and
ADR 0042's origin reservation was "stamped by the sdk, not enforced here".
Per-subject permissions make it the server's refusal.

Two things found by composing a real file rather than reading the design:

- a scoped inbox leaves a responder unable to reply, because the answer goes
  to the caller's inbox. allow_responses is the answer — one reply to the
  subject of a message actually received — and only principals that serve
  are granted it. Recorded in design 25 §4.
- composition must be deterministic: the module's entrypoint reloads on the
  file's digest, so an order-dependent composer would reload the whole bus
  on every controller restart. Covered by a test.

The golden fixture is the exact text `nats-server -t` accepts, so the syntax
is the server's rather than one we invented.
2026-09-26 20:58:49 +02:00