8e2824201a0819b1f3560f96064d9cb68599a81e
9
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8e2824201a |
Genesis can raise a mesh on the new bus, and the carried user list is checked against the composer
The mesh writes its own user list, and at genesis there is no mesh yet to write it. So the installer carries the first one — the controller's own account at a well-known bootstrap password, exactly as the store is reached at `postgres:bootstrap` and the old bus at `guest:guest`, and rotated with them. From the controller's first composition onward the file is the controller's. That left a gap I would not have found by reading: the controller's own account is created before there is a controller to mint one, so nothing recorded a hash for it, and its first composition would have left the writer out of the file it was writing — a bus nothing can connect to, produced by the thing connected to it. It now records a hash of the credential it is actually using, and only if none is recorded, so a restart cannot put the bootstrap password back over a rotated one. The carried list and the derived one are two statements of one fact, so a test compares them: every subject the controller derives must be in the template, and nothing wider. It earned itself immediately — the composer was granting both a role's whole event branch and the one event it actually follows, which is a wider way of saying the same thing, and the wider one wins. Only the submitting half of a role is granted now; what comes back is named exactly. Getting this wrong is the worst kind of silent. A controller whose carried permissions are narrower than the ones it derives comes up, connects, and is refused on the first thing it tries, with an authorisation error naming a subject and not the template that forgot it — and a mesh cannot be raised twice to find out. |
||
|
|
e4e960ec1c |
A build is work submitted to a role, on both buses
ADR 0121 carried through to working code. `Builders` is the asking side and `BuildMachine` the taking side, each with an implementation per bus, and the builder binary and the `build` command now go through them. On the bus being built, one publish does what two did. The old bus answered the asker through a reply queue and announced to an events exchange, because two audiences meant two topologies. Here the outcome is the role's own event: the asker matches it by the id its request carried, the controller records it, the catalogue places it in the graph. So a build machine publishes once, needs a reply queue for nothing, and needs a grant over nobody's inbox — which is what ruled out the alternatives. The outcome carries the module name now. Only the manifest says what was built, and on the old bus the separate announcement carried it; with one message for three readers it belongs in the result. A failed build names none, because it produced no module version and the catalogue would otherwise place something that was never made. Checked against a real server: the whole round trip; a third party on the role's event hearing the same outcome the asker did, which is the claim the decision rests on; work leaving the queue once settled, so no second machine repeats it; work submitted with no machine holding the role waiting instead of failing, and being done when one arrives; and work a machine handed back coming round again. One thing I got wrong twice now and have written down where it bit: binding to a consumer must name that consumer's own filter subject, not the narrower subject the caller cares about. The client compares the two and refuses anything that is not equal, with "subject does not match consumer". |
||
|
|
0c83ecf1b5 |
The mesh's own roles carry a protocol, and the build branch retires
ADR 0121, first half. The `mesh-*` seats said who does a job and nothing about what may be said to them or by them, so the mesh had roles it could not describe. They take the same three fields a module's seat has now, and the machinery that already derives a work queue, a holder's worker and a permission set from a declared seat does it for these too. The build-machine role accepts a build and emits an outcome, so `mesh.build.request`, `mesh.control.built` and the BUILDS stream are gone. A work queue shared by several build machines is what a seat's `accepts` already is, and keeping a second mechanism for it was two places a permission could be wrong. The controller's own side of a seat is a named list rather than something derived: it is not a module and declares no `uses`, so which roles the mesh itself submits work to has to be stated — and stating it makes that question answerable. Two things this caught: **The followed event subjects were hard-coded and had just gone stale.** They were written out while the catalogue still spelled its events as the old bus's routing keys, so converting those (issue 127) turned the pair into a controller listening to a subject nothing publishes — the same fault as the issue, from the other side. They derive from the emitter and the event name now, through the same function the permission uses, so the two cannot drift apart. **A role's queue exists before its holder**, checked against a real server, and asserting twice changes nothing. Work queues until somebody arrives to do it, so assigning a build machine later flushes the backlog instead of having lost it. |
||
|
|
d65c37caad |
The controller could not answer an enrolment, and a probe on an open server said it could
Found while reasoning about issue 127's replay question, in code committed earlier today. The controller's permissions granted no inbox at all, so the answer to every enrolment on the mesh would have been refused — "Permissions Violation for Publish to _INBOX.enrol.anchor…" — while the controller logged that it had enrolled the node. **`allow_responses` does not cover it, and that is the trap.** It permits one reply to the reply subject of a message the user received, and a message a JetStream consumer delivers has had that field claimed for the consumer's own ack address (design 25 §2). The address the controller actually answers is the one the request carried in its *payload*, which the server does not recognise as a reply subject at all. The two mechanisms look interchangeable and are not. **My earlier verification could not have caught this.** The live enrolment tests run against a server with no accounts and no permissions, so they exercise the subjects and the round trip and nothing about authority. Composing the real configuration and running a server on it is what found it. Granted the enrolment inbox space and nothing wider: nothing but an enrolling node ever subscribes under that prefix, each scoped to its own token's, so the controller publishing there is the mesh answering enrolments and reaches nothing else. Confirmed against the permissioned server both ways — the answer arrives, and a node's own inbox is still refused. Pinned as a rule that needs no server: whatever an enrolling node subscribes, the controller must be able to publish to, and a node's, a module's and a person's inbox must stay out of reach. That check is a subject-pattern match rather than a string compare, so a grant that widened by a wildcard would not slip past it. It also bears on 127's open question about who replays a build announcement: an answer to a *module's* inbox would need `_INBOX.>`, which is exactly the blanket grant design 25 §4 refuses. So the catch-up cannot become an inbox reply. |
||
|
|
f8ab9f2dcf |
The mesh composes the accounts; the module owns its server
The delivery question, decided. The alternative was a manifest field enumerating the server's ports, TLS paths and store directory so the controller could write a whole configuration file. That is wrong: those are properties of the container the module raises, they live in its image and its mounts, and the controller would have to be kept in step with a Dockerfile it never sees. So the mesh writes only what only the mesh knows — who may connect — and the module's own configuration includes it. `ComposeAccounts` is that file. A test says what must *not* be in it as plainly as what must: no port, no tls block, no store_dir. Each of those in the mesh's file is a value the controller would then own, and the module could no longer change its own image without the mesh agreeing. `bus-users` is where a module wants it written, and **asking is not enough to receive it**: the file holds every user's password hash, so a module that could ask for it could read every credential on the bus. The claim on `mesh-broker` authorises it, checked from the manifest alone. A holder with nothing composed is refused rather than given an empty file, for the reason a certificate is — a bus with no user list refuses every connection in the mesh and looks like a machine problem. Six claims checked against a running server before any of this was committed to, and two of them changed what got written: **An absolute include path is resolved relative to the including file's directory.** `include /etc/nats/accounts.conf` from /etc/nats-server/nats.conf makes the server look for /etc/nats-server/etc/nats/accounts.conf and refuse to start. So both files share one directory, and the module declares its own as a file resource beside the mesh's. **`verify: true` was refusing every connection in the mesh.** It makes the server demand a *client* certificate, and nothing in the mesh presents one: a host pins this server's exact certificate and authenticates with the password the mesh minted, and so does a module's runtime. Every connection died at the TLS handshake before any password was looked at, with an error — "client didn't provide a certificate" — that reads as a fault in the client. Removed. TLS is still required; verify only decides whether client certificates are checked. The other four: a user in an included file authenticates, an unknown user is refused so the include is the whole authority rather than an addition, a publish outside a grant is refused, and rewriting the mesh's half alone makes a new user appear — noticed by the module's own watcher, with no signal from outside, and without dropping the connection the mesh already had. That last one is task 1.2's payoff, collected. |
||
|
|
7180a273a2 |
The enrolment user is per token, and it has an inbox
Design 25 §6 says an enrolling node subscribes the inbox its own token derives. It had none: `sub` was empty, so a node would publish its request and wait out its timeout against a mesh that had answered — the handshake could not have completed. And there was one shared `enrolment` user, which cannot carry that inbox at all: a permission belongs to a user, so an inbox per token means a user per token. Named after the node, which **is** the token's id — a token is issued for a node record, the mesh holds one live claim per record, and the node's name is the one identifier both sides have before anything else is agreed. It is also exactly what the other transport does, where the account is named after the node and the secret is its password. A nameless enrolment user is now refused rather than composed into `_INBOX.enrol..>`: an empty subject token, and worse, one every nameless enrolment user would share — which is one machine able to read the credentials sealed to another. Still to wire: something that composes one of these per live token. Nothing composes enrolment users yet, on either bus — on the old one the account is made imperatively through the broker's management API when a token is issued, and here there is no management API, so issuing a token has to recompose the server's configuration. That is the remaining half of enrolment on the new bus. |
||
|
|
88bef39952 |
The consume side on NATS, and the window held by the server
The other implementation behind the seam, so the store-window guarantee now has both: one loop, one message at a time, the same window deciding. What differs is where a held message lives, and that is the whole point of the move — the AMQP side keeps an unacknowledged delivery in this process, bounded by the prefetch and lost if the controller stops; this keeps eight bytes saying when the window opened, and the message stays the server's. Checked against a running server, seven claims that reasoning cannot answer: a report is heard and leaves the work queue; one the store cannot take is naked with a delay, stays in the stream, and is recorded when the store returns; one about a superseded declaration is settled without being acted on; one the store never takes is let go once the bound passes; a heartbeat is heard and nothing is persisted; and the enrolment answer reaches the address the request carried in its payload — the test design 25 §2 asks for, so the reason for that field cannot quietly become folklore. Three things the wiring forced into the open: **The controller could not have consumed a module event.** Its permissions granted no event subject to subscribe and no ack subject on the events stream, so every announcement would have been redelivered for ever, refused by the list it already had. Both narrow: each followed subject named, not `mesh.mod.*.>`. **The controller's consumers are not derived.** It files no manifest, so its authority cannot come from a declaration that does not exist; they sit beside the mesh's own streams and are asserted the same way. No max-deliver on CONTROL — the window's bound is the controller's, and a server that dead-lettered first would discard the push the stream exists to protect. **Channels, not callbacks.** The library would run a handler on its own goroutine, and the window's bookkeeping is unlocked because the AMQP loop never had two. |
||
|
|
1f36787d75 |
The mesh's four streams, asserted on every start
Task 1.4. The foundation set only — a seat's streams come at registration and a module's consumers at assignment, neither of which has happened at genesis (ADR 0118). Asserted rather than created: a stream that was deleted, or a mesh raised from a backup, must converge rather than run without the guarantee its messages assume. Two things the definitions have to get right, both tested: - CONTROL names its subjects instead of taking mesh.control.>, because heartbeats live under that prefix and a stream of them competes for retention with the messages that matter - EVENTS filters on the event token, which is why that token exists; a filter over a module's whole namespace would persist every tool call Overlapping filters are refused where the set is written: NATS accepts two streams matching one subject and stores the message twice under two retentions, which nothing reports. Adds nats.go as a dependency; it pulled golang.org/x/* forward. Full suite green. |
||
|
|
c753f9d5c0 |
Compose the bus's accounts instead of calling a management API
Task 1.3 of novox/hq ADR 0116. On AMQP an account was an HTTP call; on NATS it is text the controller composes and the server reloads (ADR 0106). Pure, so the mesh's whole authority model is testable as strings. NATS closes a gap management.go recorded rather than hid: LavinMQ has no topic permissions, so an emitter was granted the events exchange whole and ADR 0042's origin reservation was "stamped by the sdk, not enforced here". Per-subject permissions make it the server's refusal. Two things found by composing a real file rather than reading the design: - a scoped inbox leaves a responder unable to reply, because the answer goes to the caller's inbox. allow_responses is the answer — one reply to the subject of a message actually received — and only principals that serve are granted it. Recorded in design 25 §4. - composition must be deterministic: the module's entrypoint reloads on the file's digest, so an order-dependent composer would reload the whole bus on every controller restart. Covered by a test. The golden fixture is the exact text `nats-server -t` accepts, so the syntax is the server's rather than one we invented. |