The architecture ADR 0106 asks for, proposed for review before any code: the five kinds of traffic and the guarantee each needs; one subject tree; four JetStream streams (CONTROL as a work queue carrying the store-window guarantee, NODES last-per-subject so a declaration's order is the stream's sequence — the wire-level answer to issue 107, BUILDS, EVENTS with a dead-letter stream); one NATS account with one user per module per node whose permissions are derived from emits/consumes, written as configuration the host declares and the server reloads; the nats module holding the mesh-broker seat; the enrolment handshake on an enrolment-only user; a person's client as an issued account plus a program that speaks the bus through the sdk contract (MCP as a thin adapter over it); the cutover as one rollout after the core; two lab beds. Open points listed at the end for the reviewer.
The architecture ADR 0106 asks for, **proposed for review before any code**: the five kinds of traffic and the guarantee each needs; one subject tree; four JetStream streams (CONTROL as a work queue carrying the store-window guarantee, NODES last-per-subject so a declaration's order is the stream's sequence — the wire-level answer to issue 107, BUILDS, EVENTS with a dead-letter stream); one NATS account with one user per module per node whose permissions are derived from `emits`/`consumes`, written as configuration the host declares and the server reloads; the `nats` module holding the `mesh-broker` seat; the enrolment handshake on an enrolment-only user; a person's client as an issued account plus a program that speaks the bus through the sdk contract (MCP as a thin adapter over it); the cutover as one rollout after the core; two lab beds. Open points listed at the end for the reviewer.
Also: issue 103 resolved by mesh-host #22.
Fixes, each named where it was wrong:
- reload-on is a service field; a container only has restart-on, which
recreates. Cited precedent (registry-trust-reload) is a service resource,
not a container. Fix: nats-server's own SIGHUP reload, triggered by an
in-image entrypoint watching a directory-mounted config file (issue 103's
recreate-on-change applies to a directly-mounted file, not a directory's
contents) -- asks nothing new of the host.
- A JetStream-delivered message's Reply field is already claimed by the
consumer's own ack address, so a responder using it answers nobody. Fix:
every CONTROL message needing a reply carries its reply subject in its own
payload; the controller publishes there explicitly, never via Respond().
Enrolment is the case this design actually depends on, so it's fixed there
too, not just noted.
- The listed permissions never granted publish on a durable consumer's own
ack-reply subject -- a module could receive but never ack, so every
message redelivers forever. Fixed with a scoped grant per module's own
consumer.
- One account (a deliberate choice, kept) means inbox privacy is the
permission list or nothing. The design granted 'its reply inbox' without
scoping it, which read as any user reaching any inbox. Fixed: each user's
inbox prefix is derived from its own identity and its permissions name
only that prefix.
New open question from this revision, not closed: whether the in-image
watch-and-SIGHUP shape belongs in mesh-sdk if a second module ever needs it.
Checks pass (records.py, cycle.py, index.py).
Revision addressing first review's four findings (recorded in MIGRATION-LOG.md, 2026-09-24), each fixed and named where it was wrong rather than silently patched:
reload-on/container mismatch (§5).reload-on only exists on type: service resources in mesh-host's schema (Service.ReloadOn) — the cited precedent (registry-trust-reload) is a service resource (docker.service, reloaded via systemd), not a container. A container only has restart-on, documented to mean recreate. As originally written, every account/permission change would recreate the bus server — dropped connections, lost in-flight JetStream acks, on the single resource everything else depends on. Fixed with nats-server's own native SIGHUP config reload, triggered by an in-image entrypoint watching a directory-mounted config file (issue 103's recreate-on-change compares a directly-mounted file's content, not a directory's contents) — no new host capability needed.
Eaten reply subject (§2, §6). A JetStream-delivered message's Reply field is already claimed by the consumer's own ack address by the time the controller sees it — a Respond() there answers nobody. Enrolment is the case this design actually depends on (a synchronous-feeling caller waiting on an inbox, over a subject the store-window guarantee may legitimately delay several nak cycles). Fixed: every CONTROL message needing a reply carries its reply subject as an ordinary payload field; the controller reads it from there and publishes explicitly.
Missing ack permission (§4). The listed permissions never granted publish on a durable consumer's own $JS.ACK.<stream>.<consumer>.> — a module could receive a delivery but never ack it, so it redelivers forever. Fixed with a grant scoped to each module's own consumer.
Un-scoped reply inbox under one account (§4). One account is kept (deliberate, restated why), which means inbox privacy is entirely the permission list. "Its reply inbox," ungranted a specific prefix, read as any user reaching any inbox. Fixed: each user's inbox prefix is derived from its own identity (_INBOX.<module>.<node>.>, _INBOX.person.<name>.>, _INBOX.enrol.<token-id>.>), permissions name only that prefix.
§10's bed gains coverage for all four. §11 records what's closed vs. still open, plus one new open question this revision itself raises (whether the watch-and-SIGHUP shape belongs in mesh-sdk if a second module ever needs it).
Not merged — this is a design revision for your review, same as the first version was.
Revision addressing first review's four findings (recorded in `MIGRATION-LOG.md`, 2026-09-24), each fixed and named where it was wrong rather than silently patched:
1. **`reload-on`/container mismatch (§5).** `reload-on` only exists on `type: service` resources in `mesh-host`'s schema (`Service.ReloadOn`) — the cited precedent (`registry-trust-reload`) is a service resource (`docker.service`, reloaded via systemd), not a container. A container only has `restart-on`, documented to mean recreate. As originally written, every account/permission change would recreate the bus server — dropped connections, lost in-flight JetStream acks, on the single resource everything else depends on. Fixed with `nats-server`'s own native `SIGHUP` config reload, triggered by an in-image entrypoint watching a directory-mounted config file (issue 103's recreate-on-change compares a *directly*-mounted file's content, not a directory's contents) — no new host capability needed.
2. **Eaten reply subject (§2, §6).** A JetStream-delivered message's `Reply` field is already claimed by the consumer's own ack address by the time the controller sees it — a `Respond()` there answers nobody. Enrolment is the case this design actually depends on (a synchronous-feeling caller waiting on an inbox, over a subject the store-window guarantee may legitimately delay several `nak` cycles). Fixed: every CONTROL message needing a reply carries its reply subject as an ordinary payload field; the controller reads it from there and publishes explicitly.
3. **Missing ack permission (§4).** The listed permissions never granted publish on a durable consumer's own `$JS.ACK.<stream>.<consumer>.>` — a module could receive a delivery but never ack it, so it redelivers forever. Fixed with a grant scoped to each module's own consumer.
4. **Un-scoped reply inbox under one account (§4).** One account is kept (deliberate, restated why), which means inbox privacy is entirely the permission list. "Its reply inbox," ungranted a specific prefix, read as any user reaching any inbox. Fixed: each user's inbox prefix is derived from its own identity (`_INBOX.<module>.<node>.>`, `_INBOX.person.<name>.>`, `_INBOX.enrol.<token-id>.>`), permissions name only that prefix.
§10's bed gains coverage for all four. §11 records what's closed vs. still open, plus one new open question this revision itself raises (whether the watch-and-`SIGHUP` shape belongs in `mesh-sdk` if a second module ever needs it).
Not merged — this is a design revision for your review, same as the first version was.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
The architecture ADR 0106 asks for, proposed for review before any code: the five kinds of traffic and the guarantee each needs; one subject tree; four JetStream streams (CONTROL as a work queue carrying the store-window guarantee, NODES last-per-subject so a declaration's order is the stream's sequence — the wire-level answer to issue 107, BUILDS, EVENTS with a dead-letter stream); one NATS account with one user per module per node whose permissions are derived from
emits/consumes, written as configuration the host declares and the server reloads; thenatsmodule holding themesh-brokerseat; the enrolment handshake on an enrolment-only user; a person's client as an issued account plus a program that speaks the bus through the sdk contract (MCP as a thin adapter over it); the cutover as one rollout after the core; two lab beds. Open points listed at the end for the reviewer.Also: issue 103 resolved by mesh-host #22.
Revision addressing first review's four findings (recorded in
MIGRATION-LOG.md, 2026-09-24), each fixed and named where it was wrong rather than silently patched:reload-on/container mismatch (§5).reload-ononly exists ontype: serviceresources inmesh-host's schema (Service.ReloadOn) — the cited precedent (registry-trust-reload) is a service resource (docker.service, reloaded via systemd), not a container. A container only hasrestart-on, documented to mean recreate. As originally written, every account/permission change would recreate the bus server — dropped connections, lost in-flight JetStream acks, on the single resource everything else depends on. Fixed withnats-server's own nativeSIGHUPconfig reload, triggered by an in-image entrypoint watching a directory-mounted config file (issue 103's recreate-on-change compares a directly-mounted file's content, not a directory's contents) — no new host capability needed.Replyfield is already claimed by the consumer's own ack address by the time the controller sees it — aRespond()there answers nobody. Enrolment is the case this design actually depends on (a synchronous-feeling caller waiting on an inbox, over a subject the store-window guarantee may legitimately delay severalnakcycles). Fixed: every CONTROL message needing a reply carries its reply subject as an ordinary payload field; the controller reads it from there and publishes explicitly.$JS.ACK.<stream>.<consumer>.>— a module could receive a delivery but never ack it, so it redelivers forever. Fixed with a grant scoped to each module's own consumer._INBOX.<module>.<node>.>,_INBOX.person.<name>.>,_INBOX.enrol.<token-id>.>), permissions name only that prefix.§10's bed gains coverage for all four. §11 records what's closed vs. still open, plus one new open question this revision itself raises (whether the watch-and-
SIGHUPshape belongs inmesh-sdkif a second module ever needs it).Not merged — this is a design revision for your review, same as the first version was.