Every module's event names were converted from the old bus's routing keys to local
names, and this client maps them back. If that mapping is wrong anywhere a live mesh's
events stop being delivered — silently, because a binding that matches nothing is not an
error.
So the mapping is pinned against the literal routing keys the mesh published before,
taken from the manifests as they were: what each module now emits, what each now binds,
and that a handler still matches what the bus delivers. Including the audit logger's
"everything", which must stay `#` on this bus.
And a key already in the old form is left alone, so a module built from an older manifest
keeps working beside one built from a current manifest — which is the state the mesh will
actually be in between deployments.
Design 25 §7's second item. Two surfaces over one thing — a command line for somebody
at a terminal, an MCP server for an agent — and both are adapters over the same three
calls: what tools are there, what does this one take, call it. A second way of reaching
a tool would be a second thing to keep correct.
It uses the client a module's runtime uses. Not a bridge and not a second protocol: a
person connects as their own bus user and publishes on the tool subjects their account
permits, so "what may this person do" is answered by the same permission list that
answers it for a module, and an audit has nothing separate to read.
`mesh tools` lists what the *catalogue* has, not what this credential may call. The two
differ and the difference is the point: somebody seeing only their own tools cannot tell
"not installed" from "not yours", and those need different people to fix them.
A failed call says which of three things happened, because the remedies are in three
different places: nobody serves that tool, this credential may not call it, or the tool
itself was slow. Without that they are one timeout and a stack trace.
The MCP surface decides nothing. The tool names are the ones a person types, the schemas
are the modules' own, and an answer is passed through unshaped — an adapter that
summarised somebody else's answer would be deciding what matters in it. A tool that fails
comes back as a tool error rather than a protocol error, because the request was
well-formed and the mesh answered it.
Written against the protocol directly: it is three methods and one framing, and a
dependency here would be a dependency on every workstation.
Tests drive both surfaces against a real bus, including that a host's notification is
answered with nothing and an unknown method is refused. They run one file at a time,
because each stands up a module serving the same tool subjects and run together their
requests get split between them — which showed up as one test reading another's answer.
Read back from the stream rather than from the client that wrote it, so the
check is what reached the wire. The runner lives with the implementation;
the fixture stays in one place.
Task 3.6 of novox/hq ADR 0116. A module is still written against request,
handle, publish, subscribe, close; only what is underneath changes. main.ts
still selects the AMQP client — steps 1 to 4 leave every node on AMQP, so
this ships beside it and is selected at the rollout.
Round-tripped against a real server (test/roundtrip.mjs): a tool answered
across two connections, a throwing handler reaching the caller as an error
rather than a timeout, an event delivered once with its key, body, node and
event id intact, and an event landing under its emitter's own namespace.
Three things the compiler and the server corrected:
- the envelope's field is `key`, not `type`, and the payload is `env.body`
with metadata in headers — not the whole envelope re-encoded. An
implementation that nested the envelope would pass all its own tests and
agree with nobody, which is what the conformance suite exists to stop.
- the NATS client's TLS options are PEM strings with no verify hook, so the
AMQP client's `checkServerIdentity: () => undefined` has no equivalent.
The fingerprint check still happens and is still the guarantee, but the
bus's certificate must now carry a SAN matching the address nodes dial.
That is a constraint on the mesh's certificates, recorded where it bites.
- a durable consumer is bound, never created: a module's account cannot
reach the JetStream API, and a runtime creating its own would be a module
choosing its own delivery semantics.
The retry loop treated everything but a cert-pin mismatch as transient,
so a refused login (revoked/mis-sealed credential) or a malformed broker
URL retried for ever logging 'not reachable yet' — the silent
non-progress the fix set out to remove, and a contradiction of its own
docstring. fatalBrokerReason now classifies those three as fatal and
everything else (connection refused, timeout, DNS) as retryable, with a
unit test covering the split — the honest proof the bed cannot give,
since it only ever starts the consumer after the broker is up.
The pin case is now a typed PinMismatchError caught by instanceof, not a
prose substring a reword could silently downgrade to an infinite retry
against an impostor. Added a little jitter so modules do not stampede a
recovering broker in lockstep.
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Replies go through the RPC exchange keyed by the caller's reply-queue name, not
the default exchange — so a serving module's scoped account answers with write
on mesh.rpc alone, never the default exchange (which would let it publish into
any queue). 'mesh-tools invoke <module> <tool> [args]' is the caller's side, the
sibling of emit. Verified against a real broker: a scoped account serves its
tool and is refused another module's serve queue.
The AMQP adapter now honours the full contract: events published
persistent with metadata in headers; a durable per-consumer queue
(<node>.<module>.events) with prefetch and a dead-letter exchange
(mesh.events.dead); manual ack for at-least-once.
Failure paths, not just the happy one:
- a confirm channel, so a publish the broker never accepted fails the
emit rather than vanishing — at-least-once starts at the emitter;
- a handler that keeps failing is requeued once, then dead-lettered
(poison set aside, never looping);
- an undecodable body is dead-lettered at once — it never decodes on
redelivery, and must not wedge the queue.
Binding-conformance tests against a disposable broker (a stand-in for the
mesh-hosted broker, ADR 0001): headers on the wire with a pure body, the
redelivery-limit dead-letter, and the poison-body dead-letter.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The per-node process that makes a module's tools serve on the mesh:
- broker-amqp.ts: a concrete AMQP implementation of the sdk's Broker
contract (request/reply over a reply queue + correlation id, publish/
subscribe over a topic exchange). Kept here, not in the sdk, so a
broker-client change never rebuilds a module (ADR 0044).
- runtime.ts: bind the broker, import the assigned modules' tool
entrypoints (each registers as it loads), serveTools. The thin wrapper.
- main.ts: the entrypoint, configured by MESH_BROKER_URL + MESH_TOOL_MODULES.
- Dockerfile: the container a node runs it as.
Verified over a REAL broker: the test spins LavinMQ (the mesh's broker),
the runtime serves a registered tool, a separate connection invokes it by
name over AMQP and gets the result, and an unknown tool is refused over
the wire. So 'modules can serve' is now running-on-the-mesh, not just
proven in a mock.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF