Commit Graph
401 Commits
Author SHA1 Message Date
jschoubben 0c83ecf1b5 The mesh's own roles carry a protocol, and the build branch retires
ADR 0121, first half. The `mesh-*` seats said who does a job and nothing about what
may be said to them or by them, so the mesh had roles it could not describe. They
take the same three fields a module's seat has now, and the machinery that already
derives a work queue, a holder's worker and a permission set from a declared seat
does it for these too.

The build-machine role accepts a build and emits an outcome, so `mesh.build.request`,
`mesh.control.built` and the BUILDS stream are gone. A work queue shared by several
build machines is what a seat's `accepts` already is, and keeping a second mechanism
for it was two places a permission could be wrong.

The controller's own side of a seat is a named list rather than something derived: it
is not a module and declares no `uses`, so which roles the mesh itself submits work to
has to be stated — and stating it makes that question answerable.

Two things this caught:

**The followed event subjects were hard-coded and had just gone stale.** They were
written out while the catalogue still spelled its events as the old bus's routing keys,
so converting those (issue 127) turned the pair into a controller listening to a
subject nothing publishes — the same fault as the issue, from the other side. They
derive from the emitter and the event name now, through the same function the
permission uses, so the two cannot drift apart.

**A role's queue exists before its holder**, checked against a real server, and
asserting twice changes nothing. Work queues until somebody arrives to do it, so
assigning a build machine later flushes the backlog instead of having lost it.
2026-09-27 15:34:44 +02:00
jschoubben 05ff6065d0 Event names are checked now, per manifest and across the catalogue
Issue 127 stood because nothing compared the two halves. Every manifest was
well-formed on its own and every derivation correct on its own, and no
cross-module subscription in the mesh matched anything — a subscription that
matches nothing is not an error, it is silence.

Two checks, because the mistake is possible at two scales.

Per manifest: an event is a local name, and `module.` is refused with the name to
write instead. A module emitting under what reads as another module's name is
refused too, pointing at the seat, where a name outlives whoever holds it.

Across the catalogue: where a consumed event's emitter is present, it must emit
that event. It cannot demand a live emitter for everything — a module lives in its
own repository and may be installed long before the one whose events it wants — so
the rule is narrower and still catches this. It found two real dangling
subscriptions the moment it ran.

Wildcards were undecided and two manifests needed them: `*` is one name and `**`
is the rest, spelled the mesh's way and derived to `>` here and `#` on the old bus.
A manifest naming either would stop being true when the wire changed, which is the
whole reason names are local.

And the field documentation taught the old form, examples included — which is why
the drift was uniform across 37 manifests rather than scattered. Nobody was
guessing; everybody followed the comment.
2026-09-27 14:43:16 +02:00
jschoubben d65c37caad The controller could not answer an enrolment, and a probe on an open server said it could
Found while reasoning about issue 127's replay question, in code committed earlier
today. The controller's permissions granted no inbox at all, so the answer to
every enrolment on the mesh would have been refused — "Permissions Violation for
Publish to _INBOX.enrol.anchor…" — while the controller logged that it had
enrolled the node.

**`allow_responses` does not cover it, and that is the trap.** It permits one reply
to the reply subject of a message the user received, and a message a JetStream
consumer delivers has had that field claimed for the consumer's own ack address
(design 25 §2). The address the controller actually answers is the one the request
carried in its *payload*, which the server does not recognise as a reply subject at
all. The two mechanisms look interchangeable and are not.

**My earlier verification could not have caught this.** The live enrolment tests run
against a server with no accounts and no permissions, so they exercise the subjects
and the round trip and nothing about authority. Composing the real configuration and
running a server on it is what found it.

Granted the enrolment inbox space and nothing wider: nothing but an enrolling node
ever subscribes under that prefix, each scoped to its own token's, so the controller
publishing there is the mesh answering enrolments and reaches nothing else. Confirmed
against the permissioned server both ways — the answer arrives, and a node's own
inbox is still refused.

Pinned as a rule that needs no server: whatever an enrolling node subscribes, the
controller must be able to publish to, and a node's, a module's and a person's inbox
must stay out of reach. That check is a subject-pattern match rather than a string
compare, so a grant that widened by a wildcard would not slip past it.

It also bears on 127's open question about who replays a build announcement: an
answer to a *module's* inbox would need `_INBOX.>`, which is exactly the blanket
grant design 25 §4 refuses. So the catch-up cannot become an inbox reply.
2026-09-27 14:13:52 +02:00
jschoubben eb72ec36ba 1.7 finished: minting, the file delivered, and a test flake I caused
**First, a correction: the previous commit went in on a false check.** Its message
says the suite passed; it did not. The check piped `go test` through a filter that
swallowed the failures and then printed "green" regardless. Two tests were failing
when 4de10e3 landed.

What was failing was my own doing. Purging the streams instead of deleting them
(4de10e3) left the *consumers* behind, because deleting a stream takes its
consumers with it and purging does not. A durable push consumer surviving between
tests keeps pushing to a delivery subject the previous test's subscription has gone
from: the messages count as delivered, go nowhere, and the next test waits out its
timeout for an announcement the server believes it already sent. Consumers are now
removed with the purge. Five consecutive clean runs.

`-p 1` stays, because two packages asserting and deleting the same fixed-name
objects on one bus is a real race — but its comment said the cause I had guessed
and not the one I found, so it now says the right thing.

**And delivery was not finished when I said it was.** Nothing filled
`Rendering.BusUsers`, so the composed file would never have reached a node.
`composeBusUsers` closes it: composed per push for the machine holding
`mesh-broker`, never kept, because the list is a function of the mesh's records and
a stored copy could disagree with them while both looked consistent. A user with no
credential is left out and named rather than written as a user without a password —
an ordinary situation with an obvious remedy — but a file with no users at all is
refused, because that bus would refuse every connection in the mesh.

**Minting, on both halves.** A node at enrolment and a module at `module issue`.
Three things differ from a management call and each is the point of the move: the
credential is minted into the mesh's records and becomes usable at the next
composition, so no server need be reachable; the password travels beside the address
rather than inside it, because a credential embedded in a URL leaks into every log
line that prints a connection; and a module's durable consumer is derived from what
it declared rather than named, so it cannot ask for delivery of something it did not
say it consumes.

A node reconnecting may be refused until that composition reaches the machine
running the bus. That is what the host's reconnect backoff is for and it is
survivable by design; waiting for the push would hold an enrolment open for as long
as a declaration takes to apply.

Tested that the switch is a switch: a node enrolling on one bus comes away with a
credential for that bus and none for the other, because one that held both could be
half-moved and nothing would say which half.
2026-09-27 03:19:41 +02:00
jschoubben 4de10e32e3 The bus's objects are raised on every start, and one switch says which bus
Two of 1.7's three remaining pieces.

**Raised on every start, not created once at genesis.** A stream somebody deleted,
a mesh raised from a restored backup, or a bus whose data directory was replaced
all have records and no objects — and a node whose consumer is missing hears
nothing while everything else about it looks correct.

The order is not a preference: a consumer on a stream that does not exist is
refused *naming the stream*, so somebody reading that refusal goes looking for a
deletion instead of a reversed pair of lines. Pinned by a test, along with the one
thing about seats that reads like an omission and is not — a seat's work queue is
asserted whether or not anybody holds it, because work queues until a holder
appears, so installing the module a week later flushes the backlog instead of
having lost it.

Against a real server: every object accepted, asserting twice changes nothing (a
start that failed the second time is a controller that cannot restart), a machine
joining an already-raised bus is accepted, each node's consumer is bound to its own
declaration subject and no other's, and CONTROL does not dead-letter — because the
store window's bound belongs to the controller and a server that gave up first
would discard the push the stream exists to protect.

**Which bus this mesh is on is one fact, read in one place.** Every seam the change
went behind ships both implementations; this is what the rollout flips. Being told
about both is refused at start rather than warned about: a mesh half on each is one
where a declaration goes out on one bus and the report comes back on the other, and
every component logs success while it happens — ADR 0074's failure arriving through
configuration instead of through code. The refusal names both variables and says
which to unset, because whoever reads it has to choose and the wrong choice is a
rollout half done.
2026-09-27 02:59:05 +02:00
jschoubben f8ab9f2dcf The mesh composes the accounts; the module owns its server
The delivery question, decided. The alternative was a manifest field enumerating
the server's ports, TLS paths and store directory so the controller could write a
whole configuration file. That is wrong: those are properties of the container the
module raises, they live in its image and its mounts, and the controller would
have to be kept in step with a Dockerfile it never sees. So the mesh writes only
what only the mesh knows — who may connect — and the module's own configuration
includes it.

`ComposeAccounts` is that file. A test says what must *not* be in it as plainly as
what must: no port, no tls block, no store_dir. Each of those in the mesh's file
is a value the controller would then own, and the module could no longer change
its own image without the mesh agreeing.

`bus-users` is where a module wants it written, and **asking is not enough to
receive it**: the file holds every user's password hash, so a module that could ask
for it could read every credential on the bus. The claim on `mesh-broker`
authorises it, checked from the manifest alone. A holder with nothing composed is
refused rather than given an empty file, for the reason a certificate is — a bus
with no user list refuses every connection in the mesh and looks like a machine
problem.

Six claims checked against a running server before any of this was committed to,
and two of them changed what got written:

**An absolute include path is resolved relative to the including file's
directory.** `include /etc/nats/accounts.conf` from /etc/nats-server/nats.conf
makes the server look for /etc/nats-server/etc/nats/accounts.conf and refuse to
start. So both files share one directory, and the module declares its own as a
file resource beside the mesh's.

**`verify: true` was refusing every connection in the mesh.** It makes the server
demand a *client* certificate, and nothing in the mesh presents one: a host pins
this server's exact certificate and authenticates with the password the mesh
minted, and so does a module's runtime. Every connection died at the TLS handshake
before any password was looked at, with an error — "client didn't provide a
certificate" — that reads as a fault in the client. Removed. TLS is still
required; verify only decides whether client certificates are checked.

The other four: a user in an included file authenticates, an unknown user is
refused so the include is the whole authority rather than an addition, a publish
outside a grant is refused, and rewriting the mesh's half alone makes a new user
appear — noticed by the module's own watcher, with no signal from outside, and
without dropping the connection the mesh already had. That last one is task 1.2's
payoff, collected.
2026-09-27 02:50:23 +02:00
jschoubben ee1b8ffe24 1.7, second half: the user list read out of the mesh's records
The derivation had nothing feeding it. `BusRecords` reads what it needs — the
machines, what each runs, every manifest, and which machines hold a live token —
and turns it into the records the composer derives from.

**A module's authority comes from its manifest, not from its assignment.** The
assignment says where it runs; what it may say is what it declared. So the two are
read together and the manifest decides, which is also why a seat's protocol is
gathered across the whole catalogue rather than from one manifest: a seat is
declared by one module and held by another, and that is the whole reason a seat
exists.

Three things checked against a real store, each a user that would be wrong in a
way nothing reports:

- A module assigned to a machine becomes a user with exactly the authority it
  declared, including the protocol of a seat some *other* module declared — a
  module granted nothing on a seat it was assigned to send to would fail on its
  first publish with an authorisation error that says nothing about a seat.
- Only a machine holding a live token gets an enrolment user. One outliving its
  token is a right to join that nobody issued.
- A module assigned and absent from the catalogue is refused rather than composed
  with an empty permission list. The catalogue already refuses to forget an
  assigned module, so this is the second line — and it earns its place there,
  because relying on another package's invariant is how a rule ends up enforced by
  nothing.

People are left empty rather than guessed at: the account model is built and
`operator issue` is not, so there is nobody to derive yet.
2026-09-27 01:49:09 +02:00
jschoubben 0560c792d8 1.7, first half: the mesh can say who its bus users are, and hold their keys
Two pieces the composer has been waiting for since it was written.

**The credential has to outlive its own minting.** On the bus the mesh runs on
today an account is a management call: mint a password, hand it over, seal the
plaintext to whoever will use it, keep nothing — which works because the broker
remembers. Here the users are one file, rewritten whenever any of it changes, so
keeping nothing would mean the first person's access change silently blanking
every module's password. So a bus user's bcrypt hash is now recorded, keyed by the
username the file needs, and the plaintext comes back exactly once. Verified
against a real store that the hash verifies the password it was made from, that
the password itself is not in there, that minting again rotates rather than adds,
and that forgetting a node takes its host's and its modules' credentials with it.

**Permissions are not stored, and that is the point.** Only the credential is
kept. Authority is derived from what each module declares, every time the file is
written (ADR 0043) — a stored permission list would be a second account of a
user's authority, able to disagree with the records it came from, and both would
look internally consistent while they did.

`Users` derives the list: the controller always first and always present, one
user per node, one per module per node, one per live token, one per person. Two
users with one name is refused where both can be named, rather than left to be
whichever one the server happened to read. A user the mesh has never minted a
password for is *named* rather than dropped or written as a user anybody is:
that is an ordinary situation with an obvious remedy, and the caller decides
whether a partial file is worth writing.

What remains of 1.7: delivering the file to the node that runs the server, and
minting at enrolment and assignment — which is transport-coupled, because a node
on the old bus must not be handed a credential for the new one.
2026-09-27 01:44:53 +02:00
jschoubben 7180a273a2 The enrolment user is per token, and it has an inbox
Design 25 §6 says an enrolling node subscribes the inbox its own token derives.
It had none: `sub` was empty, so a node would publish its request and wait out
its timeout against a mesh that had answered — the handshake could not have
completed.

And there was one shared `enrolment` user, which cannot carry that inbox at all:
a permission belongs to a user, so an inbox per token means a user per token.
Named after the node, which **is** the token's id — a token is issued for a node
record, the mesh holds one live claim per record, and the node's name is the one
identifier both sides have before anything else is agreed. It is also exactly
what the other transport does, where the account is named after the node and the
secret is its password.

A nameless enrolment user is now refused rather than composed into
`_INBOX.enrol..>`: an empty subject token, and worse, one every nameless
enrolment user would share — which is one machine able to read the credentials
sealed to another.

Still to wire: something that composes one of these per live token. Nothing
composes enrolment users yet, on either bus — on the old one the account is made
imperatively through the broker's management API when a token is issued, and
here there is no management API, so issuing a token has to recompose the server's
configuration. That is the remaining half of enrolment on the new bus.
2026-09-27 01:31:34 +02:00
jschoubben e65b3950cc A node's declaration consumer, which only the mesh can make
A host's account reaches no part of the JetStream API — correctly, because the
controller is the only writer of consumer definitions — so the object a node
reads its declarations through has to be waiting before the host binds to it,
and nothing created one. Named after the node, because the node's own ack grant
is `$JS.ACK.NODES.<node>.>` and a consumer named anything else is one the host
cannot acknowledge a delivery from.

No max-deliver, and a five-minute ack wait: a declaration is settled only after
the node has applied it and reported, which is minutes on a machine pulling
images, and the stream holds exactly one message per node — so there is nothing
to dead-letter, only one message to redeliver for as long as that node is away.

Asserted on start as well as created at enrolment, for the reason the streams
are: a mesh raised from a restored backup has node records and no consumers, and
a node whose consumer is missing hears nothing while everything else about it
looks correct.

The test sets the consumer against the grant the node actually gets, because
each of the three ways of getting it wrong is silent: a wrong name cannot ack, a
wrong filter reads another node's declarations, and a pull consumer is one a host
has no authority to bind.
2026-09-27 01:24:46 +02:00
jschoubben 88bef39952 The consume side on NATS, and the window held by the server
The other implementation behind the seam, so the store-window guarantee now has
both: one loop, one message at a time, the same window deciding. What differs is
where a held message lives, and that is the whole point of the move — the AMQP
side keeps an unacknowledged delivery in this process, bounded by the prefetch
and lost if the controller stops; this keeps eight bytes saying when the window
opened, and the message stays the server's.

Checked against a running server, seven claims that reasoning cannot answer: a
report is heard and leaves the work queue; one the store cannot take is naked
with a delay, stays in the stream, and is recorded when the store returns; one
about a superseded declaration is settled without being acted on; one the store
never takes is let go once the bound passes; a heartbeat is heard and nothing is
persisted; and the enrolment answer reaches the address the request carried in
its payload — the test design 25 §2 asks for, so the reason for that field
cannot quietly become folklore.

Three things the wiring forced into the open:

**The controller could not have consumed a module event.** Its permissions
granted no event subject to subscribe and no ack subject on the events stream,
so every announcement would have been redelivered for ever, refused by the list
it already had. Both narrow: each followed subject named, not `mesh.mod.*.>`.

**The controller's consumers are not derived.** It files no manifest, so its
authority cannot come from a declaration that does not exist; they sit beside the
mesh's own streams and are asserted the same way. No max-deliver on CONTROL —
the window's bound is the controller's, and a server that dead-lettered first
would discard the push the stream exists to protect.

**Channels, not callbacks.** The library would run a handler on its own
goroutine, and the window's bookkeeping is unlocked because the AMQP loop never
had two.
2026-09-27 00:53:33 +02:00
jschoubben 06cf3c04e5 The consume side behind a seam, and the window wiring into the loop
The outbound half went behind `Bus` and the transport stopped reaching its
callers; this is the other half, and the larger one. Every handler took
`amqp.Delivery`, so the serving loop could not move to another bus without
moving enrolment, reports, builds, upgrades and catch-up with it in one breath.

`Control` states one message in the mesh's words — took it, dropped it, or held
it for the store — and `Inbound` is where messages come from. The AMQP
implementation is today's loop moved rather than changed: same queues, same
prefetch, same holding, because the mesh is running on it and a bus nothing
speaks yet is no reason to alter the one every node is on.

The window (window.go) is now what decides, instead of the conditions that were
inlined in the loop. Two things that surfaced in the wiring:

**Supersession is asked before the store, not after.** A report about a
declaration the mesh has moved past would otherwise wait out a restarting store
to be written and then overwrite what the node is doing now.

**Half of a report is not about a declaration, and that half is never stale.**
What the machine *is* — the tunnel it took over, the ports its own bundle
holds, what an adopted node found, a node moving its overlay key — reaches the
mesh on a report and nowhere else. A rekey set aside as stale is a node whose
overlay key never moves, and no retry is coming, because the node said it once.
So staleness is asked only of a report that is purely an apply's account.

The one thing holding-in-memory can do that holding-in-the-server cannot is
named rather than hidden: `About` sets aside a held message when a newer one
about the same thing arrives, and the bus being built ignores it because the
digest answers the same question.
2026-09-27 00:44:16 +02:00
jschoubben 7a8a19b11b A person's account (step 4.4, the account half)
Design 25 §7. A person is not a module and holds no seat: nothing is
addressed to them, nothing is delivered to them, and they have no durable
consumer. What they have is permission to ask, as a list of tools or `*`
for an administrator.

Four properties the tests hold it to, each of which is a way of being
wrong that would not announce itself: a person reaches nothing but tools,
so one cannot claim a module said something; no ack subject, because
authority over a consumer that does not exist is authority nobody would
audit; no allow_responses, because a person who can answer a request is
impersonating a module on a bus where anyone may serve a tool; and two
people do not share an inbox.
2026-09-27 00:17:52 +02:00
jschoubben ce6ac057f6 gofmt the probe 2026-09-27 00:14:47 +02:00
jschoubben abce68fdf0 Verify that a reply address does not survive a stream
Design 25 §2 builds enrolment around carrying the reply subject in the
payload, because a JetStream consumer claims the transport Reply field for
its own ack address. The whole handshake rests on it, so it is checked:
the caller asked for _INBOX.LCr3M83q... and the consumer saw
$JS.ACK.PROBE.probe_consumer... The design was right, and the workaround
is necessary rather than defensive.

Worth having as a test rather than a note: if a future server version
stopped doing this, enrolment would keep working and the reason for the
payload field would quietly become folklore.
2026-09-27 00:14:32 +02:00
jschoubben cc019908c2 Asking a tool goes through the seam, and loses two problems
Step 3.4. On the new bus there is no reply queue to declare and no
correlation to check: each account is granted one inbox prefix and no
other, so an answer cannot reach the wrong asker. That settles a cost
build.go records having paid — on a shared reply exchange every asker saw
every result, which is why the correlation was checked rather than assumed.

And a tool nobody serves says so at once rather than after the whole wait.
The difference between "that module is down" and "that tool is slow" is the
first thing a person asking wants, and both tests are against a real server
because both are claims about what the server does, not about this code.

RequestBuild stays as it is, and is a different shape on the new bus rather
than the same one: a build takes minutes, so it is work submitted to a
queue with the outcome returning to a reply subject the request carries —
the pattern design 25 §2 already sets for anything crossing a stream. It
touches the builder too, so it goes with that conversion.
2026-09-27 00:11:30 +02:00
jschoubben d0a9abcb1c The store window as a decision, and the problem moving it to the server
introduces

Step 3.4, the consume side's hard part. The guarantee (ADR 0083) is that a
push the controller cannot record because its store is restarting is held
and retried — never dropped, never falsely acknowledged. Keeping the
delivery unacknowledged in memory becomes a nak with a delay: the server
holds it, the controller keeps no list of parked messages, and a controller
that restarts mid-window loses nothing it was holding.

That is a plain win and it introduces one problem. Holding in memory let
the controller drop an older report when a newer one for the same node
arrived, "because acting on it after the newer would undo the newer". A
naked message is the server's and comes back whatever happened meanwhile,
so the older report is redelivered after the newer was applied.

The answer was already in the message. A report carries `Declared`, the
digest of the declaration it is about, which exists because an earlier
attempt to order reports by time lost the race it invited. So supersession
stops being something the controller remembers and becomes something it
checks — the same shape as a node refusing a superseded declaration by
sequence (issue 107): ordering settled by what a message says, not by when
it arrived.

Pure, so the guarantee is testable without a bus, a store or a clock. Nine
tests, including that staleness is decided before the store is waited on —
a redelivery that lost its race must not hold a slot a current message
needs.
2026-09-27 00:05:52 +02:00
jschoubben 6c12780abe Describe the bus on its own terms
Comments framed the new bus by what it replaces — a comparison in almost
every explanation, which reads as though NATS were a variant of the old
thing rather than the mesh's nervous system. Removed throughout, and
OverAMQP becomes OverCurrent: the seam's two sides are the bus the mesh
runs on today and the one being built, not two protocols.

What remains is the client library's own package name, which is its name.
2026-09-26 23:51:00 +02:00
jschoubben 92d87b0082 gofmt bus.go — import grouping 2026-09-26 23:47:26 +02:00
jschoubben 2fad32767e The controller's outbound link behind a seam, with both transports
Step 3.4, first half. Every one of these took an *amqp.Channel, so the
transport reached every caller and swapping it meant touching all of them.
The seam turned out to be small — the controller sends exactly two kinds of
message that expect no answer — which is the same measurement that said
this bus could be replaced at all.

Bus is stated in the mesh's words, not a transport's: PublishEvent and
PublishDeclaration. Two implementations, both shipping, because steps 1 to
4 leave every node on AMQP and the NATS one is selected at the rollout.
Both ship is also what makes them comparable: one conformance fixture holds
both to the same envelope, and the NATS one is checked against a real
server reading back from the stream rather than from the code that wrote it.

Still on *amqp.Channel: RequestBuild and Ask, which carry reply-queue
machinery, and the whole consume side — the control loop, enrolment, serve.
2026-09-26 23:47:15 +02:00
jschoubben 1801f1178e Hold the Go emitter to the shared fixtures
Every required header set, each value in the pinned shape, and the subject
derived the same way. Read from the sdk's conformance directory by sibling
path, never copied.
2026-09-26 23:40:59 +02:00
jschoubben bf82805cfd A manifest holds no subject, checked across all 72
Design 29 §1's load-bearing rule: a module names its events, tools and
seats locally and the mesh derives the subject, so reorganising the subject
space leaves every manifest correct. It held by construction, and a rule
held by construction is one a later field breaks quietly.
2026-09-26 23:34:42 +02:00
jschoubben 173c8c7c21 Rename the mesh's seats to mesh-*, keeping their interfaces
novox/hq ADR 0118: the prefix is the reservation rule, so a module
declaring any mesh-* name is refused and there is no reserved-names list to
drift. Ten seats renamed in the table, the manifests that claim them, the
controller's own shipped manifests, and the tests.

Not the migration 0118 expected: a holding is derived at resolution from
manifests and never stored, so nothing recorded points at an old name. A
kept rename table tells a manifest written against one what it became —
kept rather than retired, because a module lives in its own repository and
may be registered long after the catalogue stopped using it.

**A seat is not the interface it delivers.** The git seat became mesh-git
and the git provision did not; likewise the package registry. A blanket
replace renamed both, and the failure read "the package registry is served
on <nil>", which does not say "you renamed an interface". A test now pins
every seat against what it delivers, and that neither name is also the
other.
2026-09-26 23:07:09 +02:00
jschoubben aa74bd86ca Derive a seat's stream and a module's consumer, and wire JetStream
Task 3.9's other half and 1.4's missing client. The derivation is pure and
unit-tested; only "does the server accept this" needs one running, behind
MESH_TEST_NATS so the ordinary suite stays offline.

A seat's work queue is created at registration, not assignment, so work
queues until a holder appears — a stream created at assignment would make
"the holder is not here yet" mean "your messages are gone". Named after the
seat, because the holder can change and the queued work must not care.

A holder's worker uses a queue group even though the seat guarantees one
holder: the seat is authority, the queue group is delivery, and tying them
together means the day somebody allows two holders every message is
processed twice with nothing reporting it.

One consumer per module carrying every filter, because its ack permission is
derived from its name.

And a real bug the live server caught: a durable name may not contain a dot,
but an ack subject is $JS.ACK.<stream>.<consumer>, so the single string that
read correctly inside the permission was rejected as a consumer name. Split
in two, beside the permission that has to match. Unfixed, the symptom would
have been every message redelivered forever with a permission list that
looks right — which is the failure design 25 §4 warns about.
2026-09-26 22:28:42 +02:00
jschoubben 7232d6df4b A module may declare its own seats (novox/hq ADR 0118)
The manifest carries seats and uses; registration refuses a mesh-* name, a
duplicate declarer, an undeclared uses or claim, a seat with no protocol, a
scope mismatch, and a holder that does not answer what its seat promises.

The parser stops judging unknown claim names, because it cannot: another
module may declare that seat, and one manifest cannot tell. The test that
encoded the old rule is rewritten to assert the refusal at registration, and
a new one pins the case the parser could not have distinguished.

Tools are declared for the first time, under their own key — serves already
means a provision's facts.
2026-09-26 22:17:17 +02:00
jschoubben 112cbd294d Per-subject caps on EVENTS, and a comment corrected against the server
NATS refuses overlapping streams rather than double-storing, which is the
opposite of what the Overlaps comment claimed. The check still earns its
place — it names both streams at composition rather than one at apply — and
the refusal is what rules out a shared stream beside per-module ones.
2026-09-26 21:49:55 +02:00
jschoubben 553814b6eb The mesh bus seat is one per mesh, and the amqp broker does not contend
Step 2.3 of novox/hq ADR 0116. The refusal is the resolver's existing one;
these pin it for this seat, including that a different bus implementation is
refused for the same reason — the property that lets the bus be replaced.
2026-09-26 21:40:03 +02:00
jschoubben eeb0560fc2 The mesh-broker seat delivers mesh-bus (novox/hq ADR 0120) 2026-09-26 21:17:21 +02:00
jschoubben 1f36787d75 The mesh's four streams, asserted on every start
Task 1.4. The foundation set only — a seat's streams come at registration
and a module's consumers at assignment, neither of which has happened at
genesis (ADR 0118).

Asserted rather than created: a stream that was deleted, or a mesh raised
from a backup, must converge rather than run without the guarantee its
messages assume.

Two things the definitions have to get right, both tested:
- CONTROL names its subjects instead of taking mesh.control.>, because
  heartbeats live under that prefix and a stream of them competes for
  retention with the messages that matter
- EVENTS filters on the event token, which is why that token exists; a
  filter over a module's whole namespace would persist every tool call

Overlapping filters are refused where the set is written: NATS accepts two
streams matching one subject and stores the message twice under two
retentions, which nothing reports.

Adds nats.go as a dependency; it pulled golang.org/x/* forward. Full suite
green.
2026-09-26 21:02:18 +02:00
jschoubben c753f9d5c0 Compose the bus's accounts instead of calling a management API
Task 1.3 of novox/hq ADR 0116. On AMQP an account was an HTTP call; on NATS
it is text the controller composes and the server reloads (ADR 0106). Pure,
so the mesh's whole authority model is testable as strings.

NATS closes a gap management.go recorded rather than hid: LavinMQ has no
topic permissions, so an emitter was granted the events exchange whole and
ADR 0042's origin reservation was "stamped by the sdk, not enforced here".
Per-subject permissions make it the server's refusal.

Two things found by composing a real file rather than reading the design:

- a scoped inbox leaves a responder unable to reply, because the answer goes
  to the caller's inbox. allow_responses is the answer — one reply to the
  subject of a message actually received — and only principals that serve
  are granted it. Recorded in design 25 §4.
- composition must be deterministic: the module's entrypoint reloads on the
  file's digest, so an order-dependent composer would reload the whole bus
  on every controller restart. Covered by a test.

The golden fixture is the exact text `nats-server -t` accepts, so the syntax
is the server's rather than one we invented.
2026-09-26 20:58:49 +02:00
jschoubben fb87f9f7d6 The mesh-broker seat delivers nothing
The bus is the only broker (novox/hq ADR 0117): messaging is subjects on it,
scoped by a module's own emits/consumes, not a server handed out as a
provision. So the seat joins mesh-controller and the-catalogue in delivering
no interface. Full suite green.
2026-09-26 19:34:14 +02:00
jschoubben 95426e25cf Merge pull request 'A collision is two places, not two spellings' (#71) from fix/collisions-compare-placed-paths into main 2026-09-26 16:25:40 +00:00
jschoubben fda47558d5 A collision is two places, not two spellings
checkResources compared paths as written, so ${dir:state}/server.env —
the same characters in every module, a different directory in each —
refused the first two placed modules that met. Paths are placed before
they are compared, under the default root, which keeps every real
collision: distinct modules' places are distinct under any one root,
and a module stating another's placed root is caught because a pathless
directory now owns its placed path in the comparison too.
2026-09-26 18:25:28 +02:00
jschoubben 8ea80f9584 Merge pull request 'The assignment's own root is a place, and the manifest's maps are placed' (#70) from feat/the-assignment-root-and-the-manifests-maps into main 2026-09-26 16:18:27 +00:00
jschoubben 2e3b13c0f8 The assignment's own root is a place, and the manifest's maps are placed
Slice two of ADR 0112. A pathless directory saying place "." is the
assignment's one directory, <root>/<module> — to-be 27's shape — and
place never reaches the host, which parses strictly. The maps naming
where bindings, credentials and contributions land (binds, secrets,
own-secrets, receives, grants) fill against the placed directories at
composition, into fresh maps and a fresh module slice, because one
resolution composes for many nodes. The five absolute-path checks on
those maps accept a placed reference — resolution makes it absolute
before anything reads it — while certificate, operator-keeps and
accesses paths stay absolute-only: those are the operator's or another
vocabulary's. unknownDirRefs scans the maps too, and validates place
itself: only on a directory, only ".", never beside a stated path.

Found by the foundation tests validating the sibling catalogue: the
first conversion's blanket replace turned /var/lib/gitea/database.json
into ${dir:data}base.json — which resolves to the right path by pure
string concatenation. Production was saved by a coincidence; the
catalogue cleanup that follows spells it ${dir:state}/database.json.
2026-09-26 18:18:13 +02:00
jschoubben cd2481dcd8 Merge pull request 'A directory the mesh places: ${dir:<id>} and the pathless directory resource' (#69) from feat/a-directory-the-mesh-places into main 2026-09-26 15:52:25 +00:00
jschoubben d2cbc9dbdc A directory the mesh places: ${dir:<id>} and the pathless directory resource
The first executable slice of ADR 0112 / to-be 27, sized to what the
operator settled tonight: a module definition names no host path for
its own data. A directory resource may omit path; composition resolves
it to <root>/<module>/<id>, the root a node's setting on Rendering with
/var/lib as the default — which reproduces exactly the layout novox
converged to by hand. ${dir:<id>} names the place from a resource's
path, content, mounts, environment and env-files, the same shape as
${bound:…}. A directory that states a path keeps it and still answers
by name — that is the adopted-data placement, mssql its live case.

Resolved in the controller at composition, so the wire format and the
host change not at all; a reference naming no directory refuses at the
manifest and again at composition; nested fills are rebuilt, never
written into the manifest's own maps, because one manifest composes
for many nodes.
2026-09-26 17:51:54 +02:00
jschoubben e88d3bd485 Merge pull request 'A route may say the largest body it carries' (#68) from feat/a-route-may-limit-the-body-it-carries into main 2026-09-26 14:01:56 +00:00
jochen a287812e14 A route may say the largest body it carries
Proxy configuration beside insecure, not a fifth policy — ADR 0108 closed that set at four, and both
of these tune how a request is carried rather than deciding what a name admits. A registry is the
case that needs it: image layers arrive as single requests of gigabytes and a proxy's own default
refuses them long before the workload is reached.

Absent is no limit, which is what every route already got. A limit that is not a whole positive
number of bytes takes the route with it, named in the log like a port that is not one — serving it
without the limit would carry exactly what the module said not to carry. Enforced on the declared
length where there is one, and while reading for a chunked body, which declares none: without the
second, a limit is advice.
2026-09-26 16:01:21 +02:00
jschoubben 557f419e71 Merge pull request 'mount the broker TLS directory bind, not the old named volume' (#54) from fix/own-broker-tls-mount-is-the-directory-bind into main 2026-09-26 12:55:32 +00:00
jochen 71c8080359 Declare the broker's TLS directory as an access, not an undeclared bind
The swap from a named volume to the host directory left the mount undeclared, which main's own
manifest check now refuses: a bind the module did not declare is created by the runtime as root, so
the module's owner and mode never reach it and ADR 0030's data rule does not cover it.

The directory is the broker's — lavinmq declares it as its own, mode 0700 — so from here it is an
access, read-only: a pre-existing path this module is granted use of and does not own.
2026-09-26 14:55:04 +02:00
jschoubben cc86f8b433 Merge main 2026-09-26 14:53:55 +02:00
jschoubben 0944311f86 Merge pull request 'Seats are a closed set, a seat's holder answers for what it delivers, and a build source may live on the git seat' (#63) from feat/seats-are-a-closed-set into main 2026-09-26 12:31:16 +00:00
jschoubben 7e42380dcd Merge main 2026-09-26 14:29:00 +02:00
jschoubben 41de152739 Merge pull request 'route-proxy: a policy refusal is also not-my-token' (#67) from fix/a-policy-refusal-is-also-not-my-token into main 2026-09-26 12:26:19 +00:00
jschoubben dad0a153ff route-proxy: a policy refusal is also not-my-token
autocert checks the host policy before the token and answers 403 — the
internal authority does this for every public name, so mail.novox.be's
challenge died on the internal manager's probe one commit after it
stopped dying on the public one's 404. Both shapes of refusal now fall
through to routing; a fifth test pins the 403 case with a refusing
policy.
2026-09-26 14:26:01 +02:00
jschoubben 345722a4fd Merge pull request 'route-proxy: the challenge path falls through for real' (#66) from fix/the-challenge-path-falls-through-for-real into main 2026-09-26 12:22:21 +00:00
jschoubben f145d17fc8 route-proxy: the challenge path falls through for real
autocert's HTTPHandler answers 404 itself for a token it does not hold
and never consults its fallback on the challenge path — the
predecessor's exact fault, rediscovered live when Mailu's renewal died
behind this proxy on cutover day. tokenOrRoute probes each authority
against a buffered writer and hands a token none of them holds to plain
routing, so a consumer's own ACME client answers its own challenge
through an ordinary path-scoped route. Four tests pin it, including the
cache-key shape a restart-surviving token actually has.
2026-09-26 14:21:52 +02:00
jschoubben c3458a2546 Merge pull request 'route-proxy: a second authority for internal names, and https targets' (#64) from feat/route-proxy-internal-acme into main 2026-09-25 19:54:51 +00:00
jschoubben f8919e2079 Merge pull request 'builder: a clone may offer the forge's credential, through git's own store' (#65) from feat/builder-clones-with-the-forges-credential into main 2026-09-25 19:54:39 +00:00