Task 1.3 of novox/hq ADR 0116. On AMQP an account was an HTTP call; on NATS
it is text the controller composes and the server reloads (ADR 0106). Pure,
so the mesh's whole authority model is testable as strings.
NATS closes a gap management.go recorded rather than hid: LavinMQ has no
topic permissions, so an emitter was granted the events exchange whole and
ADR 0042's origin reservation was "stamped by the sdk, not enforced here".
Per-subject permissions make it the server's refusal.
Two things found by composing a real file rather than reading the design:
- a scoped inbox leaves a responder unable to reply, because the answer goes
to the caller's inbox. allow_responses is the answer — one reply to the
subject of a message actually received — and only principals that serve
are granted it. Recorded in design 25 §4.
- composition must be deterministic: the module's entrypoint reloads on the
file's digest, so an order-dependent composer would reload the whole bus
on every controller restart. Covered by a test.
The golden fixture is the exact text `nats-server -t` accepts, so the syntax
is the server's rather than one we invented.
Three readers did not follow a moved foundation port (novox/hq 04-ISSUES/102),
and each took the control-node down in its own way: the control plane's own
store and broker connections, sealed at genesis with the port inside; and every
build the mesh ever recorded, kept as `<registry>:<port>/<module>/<artifact>@…`.
The control plane cannot open its own sealed connections to move a port, and it
cannot bind the store as a consumer would — a binding mints a credential. So its
settings get a third twin, `NAME_PORT`, read on top of the sealed value by the
store, the broker, the management API and the bus connection, and filled into
its container by a placeholder that names a seat, `${seat:mesh-store:5432}`,
from the node's given or mesh-assigned ports — never the manifest's number, and
empty when the mesh has nothing to add, so what genesis wrote stands. A value
that is still a placeholder is nothing said, aloud: the manifest naming it lands
in the next commit, once every control plane that composes it knows it.
A build is now recorded by digest and path — `artifact-store://<module>/<artifact>@…`
— and the store's address is composed in where a reference is used: the
declaration, the trust file, the bases a build is handed, a replay to the
catalogue. Over the network as `<node>.internal:<port>`; on the store's own node
before any network exists — every genesis push before its "network" step — by
loopback. A reference recorded before this, with an address, is re-routed the
same way when the mesh built it. The trust file and every provider's address
come from one derivation: the node's given port, over the mesh's assignment,
over the manifest's number.
novox/hq 04-ISSUES/102
The broker settings take a _FILE twin like the store connections; the
catalogue engine refuses a secret placeholder in a container's env and a
secret-carrying env-file unless the container says why with
secrets-in-environment, which stays in the catalogue and never reaches the
machine.
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Answering and announcing are different acts. The reply goes to whoever asked and
is correlated to their request; the announcement says to the whole mesh that a
module now exists at a commit, which is what the catalogue places in the module
graph (novox/hq ADR 0072). A build nobody asked for still has to be announced, or
the graph knows less than the registry does.
What it was built on top of is read out of the build's own inputs rather than
declared, because a declared list drifts from what the code actually uses
(ADR 0009). These are artifact references, which is what a build input names;
resolving them to module-versions is the catalogue's work, since it is what knows
which module-version published which artifact.
Events ride the topic exchange, not the direct one nodes speak over, so the
builder's account is granted both: it must be able to answer and to announce.
The envelope is the sdk's, reproduced exactly — a second shape would be a second
thing for consumers to handle, and they are written against the first.
Announcing is not allowed to fail a build. The work was done and was answered; a
build reported as failed because saying so failed is a lie about it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
CreateModuleAccount now also grants serve.<module>.* (declare, bind, consume its
own tool queues) and mesh.rpc (bind them on, publish replies) — so a module can
serve its tools and reply, scoped to exactly its own, and no other module's. The
broker tests still hold a module out of another's queue.
LavinMQ refuses a non-administrator declaring a queue with a dead-letter
exchange, so a scoped module cannot make its own. EnsureModuleQueue declares
<node>.<module>.events with its DLX as the mesh, and 'module issue' does so
for a consuming module — the runtime then passively checks it rather than
declaring. Verified against a real broker: the scoped account binds and
consumes the pre-declared queue.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
CreateModuleAccount gives an assigned module its own broker account whose
permissions ARE its manifest: declare and read its own <node>.<module>.events
queue, read the events exchange to bind onto if it consumes, write the events
exchange only if it emits. The account name carries the node (sealed per
machine), the permissions carry the module (one cannot read another's queue).
The builder becomes one instance of this rule rather than a separate kind.
EnsureEventExchanges declares the bus the substrate owns — mesh.events,
mesh.rpc, mesh.events.dead + a retention queue — idempotently, since a module
account may not declare an exchange.
Scope tested as patterns (no broker needed), and every management call verified
against a real LavinMQ. Honest limit recorded in the code: LavinMQ has no topic
permissions, so ADR 0047's emit-origin reservation (module.<self>.*) is stamped
by the sdk, not enforced by the broker; a pure consumer like the audit logger
is unaffected.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The builder was documented as holding its own broker credential and
nothing else, and nothing issued one — so in practice it used whatever it
was handed, which was the broker's administrative account. A program
documented as holding its own credential and given somebody else's is
worse than one with no story at all.
`builder issue <name>` creates an account that may read the build queue
and write to the mesh exchange. Not a node account: a build machine is
not a node, and a node's queue carries its declarations.
Two faults found by running it, both about the answer path:
- the reply queue was left for the broker to name, and the account was
scoped to `amq.gen-*` — one broker's convention. The builder built,
could not answer, and the connection closed. Reply queues are named
here now, deterministically.
- the answer then went via the DEFAULT exchange, where permission is
granted per exchange rather than per queue. A builder allowed to use it
could publish into any node's queue, which is the privilege a build
machine most obviously should not have. Answers go through the mesh
exchange, which it already may use, and an asker binds its reply queue
to the same key and filters by correlation.
Verified against a real broker: a builder cannot consume a node's queue
and cannot publish to the default exchange. That check nearly reported
the opposite — an unconfirmed publish is asynchronous, so the refusal
arrives as a channel close afterwards and a naive test sees success. With
publisher confirms it is immediate. A negative security assertion made
against an asynchronous call is not an assertion.
Redelivery was observed working while fixing this: builders that died
before answering left their work on the queue, and the next builder did
all of it.
Also: the queue and exchange names exist in both `broker` and `link`,
because `link` imports `broker`. A test in an external package keeps them
agreeing — a builder scoped to a queue nothing publishes to takes no work
and says nothing about why.
There was no chicken-and-egg to solve. The mesh runs the broker, so it creates
the node's account when it issues the token, and the one-time secret is that
account's password. A joining node's first connection is already authenticated;
enrolment is what it says once it is in. I had been treating this as a decision
that needed taking, and it did not.
The account is per node and scoped: it may read its own queue, write to the one
exchange, and configure nothing else. The patterns are anchored and the node
name is constrained to characters that cannot widen them, because a name
carrying a dot or a star would silently let that node read everybody's queues.
`serve` is the control plane running: one connection, one queue, one consumer.
One deliberately -- two consumers on a queue get round-robined and each receives
half of what it expects, which has happened on this project before, between a
module's daemon and its capability server.
Enrolment spends the token first, in the single statement that both finds and
marks it, and only then records the key. That order is the order things become
irreversible: recording a key for a node whose token turned out to be spent
would leave the mesh believing a machine that never had the right to join.
Refusals are one message for every reason. The log says which, where an
operator can see it; the node is told only that the token cannot be used.
Verified in the lab, on a sealed machine, through the whole first-node path.
Given the broker's address and its certificate, mesh-control now issues a
token carrying everything ADR 0004 asks for: where to connect, what to expect
there, whose signature to believe afterwards, and a one-time right to join.
Verified by decoding one and checking the fingerprint against `openssl x509 |
sha256sum` -- they match.
The fingerprint is derived from the certificate on disk and never configured.
A configured pin can drift from the certificate it describes, and a drifted pin
is worse than none: every node issued a token during the drift refuses to
connect, and the failure looks like an attack rather than a mistake.
Computed over DER, which is what a client sees on the wire. Hashing the PEM
text instead would mean the same certificate, re-wrapped with different line
endings, produced a different pin -- there is a test for exactly that, and one
for pointing this at tls.key by mistake, which would otherwise produce a
confident pin over the wrong file.
Having no broker stays a state rather than a failure: a control plane holds
records and a signing key without one. Having half a broker is refused, because
a token with an address and nothing to check it against invites a node to trust
whatever answers.
Fault injection caught the same weak test I wrote earlier in the day -- asking
whether something failed rather than why, so deleting the guard changed nothing
because it failed one line later anyway. Both are now asserted on the reason.