A merge-check.sh, the repository's layer of the mesh's merge check (mesh/repo-check):
format, vet, and node-tools' suite under the race detector. The tests shared one bus and
had to run one package at a time; like the controller's (#97), each now starts a server
of its own at the release go.mod pins, held to the catalogue's bus image by a test. What
cannot run in the check — bundles that need @novox/mesh-sdk — is said as not tested.
"plex did not take radarr.download.completed: Unexpected end of JSON input"
read as the runtime failing to parse the bundle's answer. It was plex's own
handler error, relayed. A bundle's error answer is now a launch.Refused, and
the line says the handler answered an error, quotes it, names the event id
and the delay before the next offer; the runtime's own failures (no answer
in time, bundle exited) are said as before.
The rule, unchanged and now tested: any mesh/event answer that is not an
error takes the event, whatever its result says, including none.
A machine whose node-engine is heard and whose runtime is gone is a
machine nobody can ask anything, and nothing said so. The runtime now
says on mesh.control.<node>.tools-alive, every minute, that it is there,
with its interval and build; the controller raises tools-silent after
three missed. Core NATS, like the host's heartbeat: a lost one is the
next one. A heartbeat the bus refuses is logged when that starts and
when it stops.
The console took node out of every schema and every call, so the
controller's push, plan, assign, pin and settings were described without
the machine and called without it: a push naming one machine ran as a
push of every machine behind (hq issue 244). node is now taken out only
where the address names the machine, a different machine there is
refused, and an argument a seat's verb does not declare is refused.
The console gathered $SRV.INFO answers for a fixed 750 ms. The laptop's
runtime (48 modules, 341 endpoints, 164 kB) answers last every time: its
answer crosses to the broker on another machine and back, a median of
365 ms on a quiet link and 813 ms in one of 25 rounds measured. When it
missed the window the console said its modules ran nowhere ("nothing in
the mesh is called slack") or only on another machine.
Discovery now asks $SRV.PING alongside $SRV.INFO and waits, past the
window and up to 5 s, for every instance that answered PING. A runtime
that never sends what it serves, or answered before and not now, is
named in every answer that might concern it instead of the module being
called missing. An answer still too large after first-line descriptions
drops them, and says so in its metadata.
mesh_runtimes reports, per runtime, its machine, how long its answer
took, its size, its modules and tools, whether it was shortened and when
it was last heard, and the runtimes and machines not heard.
Every tool's schema in one $SRV.INFO reply outgrew max_payload on machines
serving 137-206 tools, and the refused reply was dropped silently, so search
and a machine's view went empty. Module tools now announce without schema;
describe and tools/list ask the module for it. An answer still too large has
its descriptions cut to a line, and a failed reply is logged.
"authorization" alone matched a tool's own answer mentioning the word — the
runtime refusing a state value with an Authorization header read as
"this account may not call".
mesh/state.get, put, delete, keys and watch on the stdio channel, from the
buckets the membership issues. A watch hands the current values without
deletions, then every change, and is answered once the current values are
delivered. Refused with the reason: state not issued, a reader's write, a
value with a credential-named field — the bus alone would answer a refused
write with a timeout.
A push sends the machine that needs a grant and the bus's machine in the same breath, and the
runtime can subscribe the moment before the bus reloads its user list. Refused once, the
subscription stayed dead until some later membership re-served it, and a newly assigned module ran
unreachable. A refused subject the runtime answers is now asked for again for about five minutes.
The controller announces the mesh-controller seat without a machine — the seat is the mesh's — and
the console then reported it as not answering on the machine it is assigned to. An announcement that
names no machine now answers for wherever its module is assigned.
Both runtimes subscribed $SRV.<verb>.> as a wildcard; the grants allow the bare question and the
service's own name and instance. The bus refused the wildcard, and the TypeScript runtime treats a
refused subscription as fatal, so every per-module container crash-looped after the image rolled.
They now subscribe exactly $SRV.<verb>, $SRV.<verb>.<name> and $SRV.<verb>.<name>.<id>.
A launched bundle's mesh/subscribe binds the module's own durable consumer — EVENTS, <node>_<module>,
by name as the module's own runtime bound it, so nothing is lost or replayed in the move — and every
event goes to each child of the module that subscribed as mesh/event, acknowledged only when all
answered, negatively acknowledged after a short delay when one failed or died, terminated when it is
not an event. mesh/ask calls a tool as the module. Every launched bundle is started again when it
exits, with backoff, since long-running code waits for no call. Requires SDK 0.1.6.
The Go runtime answers $SRV.PING, $SRV.INFO and $SRV.STATS (and per name and id) in the
io.nats.micro.v1 format with what it serves at the moment it is asked: one service per runtime
process, since the bus admits one reply per request from each responder, and one endpoint per tool
per subject, its metadata saying module, seat, scope, machine, description, schema and whether the
module is interchangeable. Serving is unchanged.
The console gathers one $SRV.INFO request's answers instead of asking the catalogue's roster and
each module's tools, and reads the controller's records as JSON for what should have answered: an
assignment with tools that did not announce is named, a module without tools never is. The text
parsers of node list and module list are gone. Packages share the test bus: go test -p 1.
mesh_overview, mesh_machine, mesh_search, mesh_describe and mesh_call walk the mesh's structure;
every tool has one address per layer: <seat>.<verb>, <node>/<seat>.<verb>, <node>/<module>.<tool>,
and <module>.<tool> for a module the mesh issued a plain subject. A stateful module called without
its machine, a node seat without one, a mesh seat with one, or a module on the wrong machine is
refused naming what would work. Answers come from the mesh when asked, kept five seconds, so a tool
that arrives mid-session is found. The flat catalogue stays behind MESH_CONSOLE_FLAT=1 and old
<module>.<tool> names still answer.
The runtime knows no language, so nothing ties it to Node.js. This ports its serve mode — the
pinned bus connection and patient connect, following memberships, launching every served bundle
over MCP on stdio with its own environment, a child's emit published as its module, each tool,
the tools verb and seat verbs served where the mesh issued them, and the console on loopback —
to one static binary. Same subjects, request and reply bodies, event headers and MCP answers.
The TypeScript stays: it is still the runtime inside the per-module containers until WP4c.
Tests run against a real bus and share the TypeScript fixtures.