Commit Graph
474 Commits
Author SHA1 Message Date
jschoubben 05ff6065d0 Event names are checked now, per manifest and across the catalogue
Issue 127 stood because nothing compared the two halves. Every manifest was
well-formed on its own and every derivation correct on its own, and no
cross-module subscription in the mesh matched anything — a subscription that
matches nothing is not an error, it is silence.

Two checks, because the mistake is possible at two scales.

Per manifest: an event is a local name, and `module.` is refused with the name to
write instead. A module emitting under what reads as another module's name is
refused too, pointing at the seat, where a name outlives whoever holds it.

Across the catalogue: where a consumed event's emitter is present, it must emit
that event. It cannot demand a live emitter for everything — a module lives in its
own repository and may be installed long before the one whose events it wants — so
the rule is narrower and still catches this. It found two real dangling
subscriptions the moment it ran.

Wildcards were undecided and two manifests needed them: `*` is one name and `**`
is the rest, spelled the mesh's way and derived to `>` here and `#` on the old bus.
A manifest naming either would stop being true when the wire changed, which is the
whole reason names are local.

And the field documentation taught the old form, examples included — which is why
the drift was uniform across 37 manifests rather than scattered. Nobody was
guessing; everybody followed the comment.
2026-09-27 14:43:16 +02:00
jschoubben a7df0fc62f Merge pull request 'Name system seats by scope; let a module define its own (ADR 0121)' (#81) from feat/system-seats-named-by-scope into main 2026-09-27 12:32:27 +00:00
jschoubben 1c56210530 Name system seats by scope; let a module define its own (ADR 0121)
System seats are mesh-* (one, mesh-wide) or node-* (one per node). Renamed:
the-build-machine -> mesh-build-machine (+scope mesh), the-catalogue ->
mesh-catalog, the-dns-port -> node-dns-resolver, the-intrusion-prevention ->
node-intrusion-prevention, the-packet-filter -> node-packet-filter,
the-resolver-configuration -> node-resolver-config, the-uplink -> node-uplink.
Removed the-showcase from the set — it becomes the first module-defined seat.

A manifest may declare its own seats (DefinesSeats); a claim is a system seat,
a reserved mesh-*/node-* name the mesh does not define (refused), or a
module-defined seat valid only when the manifest declares it.

Deferred: the delivering registry seats (git, npm-package-registry,
the-artifact-store) and the-private-network (a scope + server/client model
change), per ADR 0121.
2026-09-27 14:30:56 +02:00
jschoubben d65c37caad The controller could not answer an enrolment, and a probe on an open server said it could
Found while reasoning about issue 127's replay question, in code committed earlier
today. The controller's permissions granted no inbox at all, so the answer to
every enrolment on the mesh would have been refused — "Permissions Violation for
Publish to _INBOX.enrol.anchor…" — while the controller logged that it had
enrolled the node.

**`allow_responses` does not cover it, and that is the trap.** It permits one reply
to the reply subject of a message the user received, and a message a JetStream
consumer delivers has had that field claimed for the consumer's own ack address
(design 25 §2). The address the controller actually answers is the one the request
carried in its *payload*, which the server does not recognise as a reply subject at
all. The two mechanisms look interchangeable and are not.

**My earlier verification could not have caught this.** The live enrolment tests run
against a server with no accounts and no permissions, so they exercise the subjects
and the round trip and nothing about authority. Composing the real configuration and
running a server on it is what found it.

Granted the enrolment inbox space and nothing wider: nothing but an enrolling node
ever subscribes under that prefix, each scoped to its own token's, so the controller
publishing there is the mesh answering enrolments and reaches nothing else. Confirmed
against the permissioned server both ways — the answer arrives, and a node's own
inbox is still refused.

Pinned as a rule that needs no server: whatever an enrolling node subscribes, the
controller must be able to publish to, and a node's, a module's and a person's inbox
must stay out of reach. That check is a subject-pattern match rather than a string
compare, so a grant that widened by a wildcard would not slip past it.

It also bears on 127's open question about who replays a build announcement: an
answer to a *module's* inbox would need `_INBOX.>`, which is exactly the blanket
grant design 25 §4 refuses. So the catch-up cannot become an inbox reply.
2026-09-27 14:13:52 +02:00
jschoubben eb72ec36ba 1.7 finished: minting, the file delivered, and a test flake I caused
**First, a correction: the previous commit went in on a false check.** Its message
says the suite passed; it did not. The check piped `go test` through a filter that
swallowed the failures and then printed "green" regardless. Two tests were failing
when 4de10e3 landed.

What was failing was my own doing. Purging the streams instead of deleting them
(4de10e3) left the *consumers* behind, because deleting a stream takes its
consumers with it and purging does not. A durable push consumer surviving between
tests keeps pushing to a delivery subject the previous test's subscription has gone
from: the messages count as delivered, go nowhere, and the next test waits out its
timeout for an announcement the server believes it already sent. Consumers are now
removed with the purge. Five consecutive clean runs.

`-p 1` stays, because two packages asserting and deleting the same fixed-name
objects on one bus is a real race — but its comment said the cause I had guessed
and not the one I found, so it now says the right thing.

**And delivery was not finished when I said it was.** Nothing filled
`Rendering.BusUsers`, so the composed file would never have reached a node.
`composeBusUsers` closes it: composed per push for the machine holding
`mesh-broker`, never kept, because the list is a function of the mesh's records and
a stored copy could disagree with them while both looked consistent. A user with no
credential is left out and named rather than written as a user without a password —
an ordinary situation with an obvious remedy — but a file with no users at all is
refused, because that bus would refuse every connection in the mesh.

**Minting, on both halves.** A node at enrolment and a module at `module issue`.
Three things differ from a management call and each is the point of the move: the
credential is minted into the mesh's records and becomes usable at the next
composition, so no server need be reachable; the password travels beside the address
rather than inside it, because a credential embedded in a URL leaks into every log
line that prints a connection; and a module's durable consumer is derived from what
it declared rather than named, so it cannot ask for delivery of something it did not
say it consumes.

A node reconnecting may be refused until that composition reaches the machine
running the bus. That is what the host's reconnect backoff is for and it is
survivable by design; waiting for the push would hold an enrolment open for as long
as a declaration takes to apply.

Tested that the switch is a switch: a node enrolling on one bus comes away with a
credential for that bus and none for the other, because one that held both could be
half-moved and nothing would say which half.
2026-09-27 03:19:41 +02:00
jschoubben 4de10e32e3 The bus's objects are raised on every start, and one switch says which bus
Two of 1.7's three remaining pieces.

**Raised on every start, not created once at genesis.** A stream somebody deleted,
a mesh raised from a restored backup, or a bus whose data directory was replaced
all have records and no objects — and a node whose consumer is missing hears
nothing while everything else about it looks correct.

The order is not a preference: a consumer on a stream that does not exist is
refused *naming the stream*, so somebody reading that refusal goes looking for a
deletion instead of a reversed pair of lines. Pinned by a test, along with the one
thing about seats that reads like an omission and is not — a seat's work queue is
asserted whether or not anybody holds it, because work queues until a holder
appears, so installing the module a week later flushes the backlog instead of
having lost it.

Against a real server: every object accepted, asserting twice changes nothing (a
start that failed the second time is a controller that cannot restart), a machine
joining an already-raised bus is accepted, each node's consumer is bound to its own
declaration subject and no other's, and CONTROL does not dead-letter — because the
store window's bound belongs to the controller and a server that gave up first
would discard the push the stream exists to protect.

**Which bus this mesh is on is one fact, read in one place.** Every seam the change
went behind ships both implementations; this is what the rollout flips. Being told
about both is refused at start rather than warned about: a mesh half on each is one
where a declaration goes out on one bus and the report comes back on the other, and
every component logs success while it happens — ADR 0074's failure arriving through
configuration instead of through code. The refusal names both variables and says
which to unset, because whoever reads it has to choose and the wrong choice is a
rollout half done.
2026-09-27 02:59:05 +02:00
jschoubben f8ab9f2dcf The mesh composes the accounts; the module owns its server
The delivery question, decided. The alternative was a manifest field enumerating
the server's ports, TLS paths and store directory so the controller could write a
whole configuration file. That is wrong: those are properties of the container the
module raises, they live in its image and its mounts, and the controller would
have to be kept in step with a Dockerfile it never sees. So the mesh writes only
what only the mesh knows — who may connect — and the module's own configuration
includes it.

`ComposeAccounts` is that file. A test says what must *not* be in it as plainly as
what must: no port, no tls block, no store_dir. Each of those in the mesh's file
is a value the controller would then own, and the module could no longer change
its own image without the mesh agreeing.

`bus-users` is where a module wants it written, and **asking is not enough to
receive it**: the file holds every user's password hash, so a module that could ask
for it could read every credential on the bus. The claim on `mesh-broker`
authorises it, checked from the manifest alone. A holder with nothing composed is
refused rather than given an empty file, for the reason a certificate is — a bus
with no user list refuses every connection in the mesh and looks like a machine
problem.

Six claims checked against a running server before any of this was committed to,
and two of them changed what got written:

**An absolute include path is resolved relative to the including file's
directory.** `include /etc/nats/accounts.conf` from /etc/nats-server/nats.conf
makes the server look for /etc/nats-server/etc/nats/accounts.conf and refuse to
start. So both files share one directory, and the module declares its own as a
file resource beside the mesh's.

**`verify: true` was refusing every connection in the mesh.** It makes the server
demand a *client* certificate, and nothing in the mesh presents one: a host pins
this server's exact certificate and authenticates with the password the mesh
minted, and so does a module's runtime. Every connection died at the TLS handshake
before any password was looked at, with an error — "client didn't provide a
certificate" — that reads as a fault in the client. Removed. TLS is still
required; verify only decides whether client certificates are checked.

The other four: a user in an included file authenticates, an unknown user is
refused so the include is the whole authority rather than an addition, a publish
outside a grant is refused, and rewriting the mesh's half alone makes a new user
appear — noticed by the module's own watcher, with no signal from outside, and
without dropping the connection the mesh already had. That last one is task 1.2's
payoff, collected.
2026-09-27 02:50:23 +02:00
jschoubben ccf50b1269 Merge pull request 'Facts carry the format as a template, so the control plane holds none (ADR 0120)' (#80) from feat/roster-facts-are-templates into main 2026-09-26 23:51:23 +00:00
jschoubben a6d89e3f73 Reconcile with hq 128: hosts template is the region form, node-names is shared
The merge commit took only the staged index; these reconciliation edits sat
unstaged in the working tree. Integrate the template mechanism with #79's
region write (hq 128): RosterFile gains Shared, FactsInto sets into:block for
a shared fact, /etc/hosts becomes the region form (no floor) and node-names is
marked shared. Without this the merge would have regressed /etc/hosts back to
a whole-file write, replacing the operator's own lines.
2026-09-27 01:50:11 +02:00
jschoubben ee1b8ffe24 1.7, second half: the user list read out of the mesh's records
The derivation had nothing feeding it. `BusRecords` reads what it needs — the
machines, what each runs, every manifest, and which machines hold a live token —
and turns it into the records the composer derives from.

**A module's authority comes from its manifest, not from its assignment.** The
assignment says where it runs; what it may say is what it declared. So the two are
read together and the manifest decides, which is also why a seat's protocol is
gathered across the whole catalogue rather than from one manifest: a seat is
declared by one module and held by another, and that is the whole reason a seat
exists.

Three things checked against a real store, each a user that would be wrong in a
way nothing reports:

- A module assigned to a machine becomes a user with exactly the authority it
  declared, including the protocol of a seat some *other* module declared — a
  module granted nothing on a seat it was assigned to send to would fail on its
  first publish with an authorisation error that says nothing about a seat.
- Only a machine holding a live token gets an enrolment user. One outliving its
  token is a right to join that nobody issued.
- A module assigned and absent from the catalogue is refused rather than composed
  with an empty permission list. The catalogue already refuses to forget an
  assigned module, so this is the second line — and it earns its place there,
  because relying on another package's invariant is how a rule ends up enforced by
  nothing.

People are left empty rather than guessed at: the account model is built and
`operator issue` is not, so there is nobody to derive yet.
2026-09-27 01:49:09 +02:00
jschoubben 0560c792d8 1.7, first half: the mesh can say who its bus users are, and hold their keys
Two pieces the composer has been waiting for since it was written.

**The credential has to outlive its own minting.** On the bus the mesh runs on
today an account is a management call: mint a password, hand it over, seal the
plaintext to whoever will use it, keep nothing — which works because the broker
remembers. Here the users are one file, rewritten whenever any of it changes, so
keeping nothing would mean the first person's access change silently blanking
every module's password. So a bus user's bcrypt hash is now recorded, keyed by the
username the file needs, and the plaintext comes back exactly once. Verified
against a real store that the hash verifies the password it was made from, that
the password itself is not in there, that minting again rotates rather than adds,
and that forgetting a node takes its host's and its modules' credentials with it.

**Permissions are not stored, and that is the point.** Only the credential is
kept. Authority is derived from what each module declares, every time the file is
written (ADR 0043) — a stored permission list would be a second account of a
user's authority, able to disagree with the records it came from, and both would
look internally consistent while they did.

`Users` derives the list: the controller always first and always present, one
user per node, one per module per node, one per live token, one per person. Two
users with one name is refused where both can be named, rather than left to be
whichever one the server happened to read. A user the mesh has never minted a
password for is *named* rather than dropped or written as a user anybody is:
that is an ordinary situation with an obvious remedy, and the caller decides
whether a partial file is worth writing.

What remains of 1.7: delivering the file to the node that runs the server, and
minting at enrolment and assignment — which is transport-coupled, because a node
on the old bus must not be handed a credential for the new one.
2026-09-27 01:44:53 +02:00
jschoubben 966c5ddad3 Merge remote-tracking branch 'origin/main' into feat/roster-facts-are-templates
# Conflicts:
#	internal/catalogue/facts.go
#	internal/catalogue/facts_test.go
2026-09-27 01:41:23 +02:00
jschoubben b36f822cb6 Facts carry the format as a template, so the control plane holds none (ADR 0120)
A roster fact used to be a name from a closed list, each formatted in Go
here — node-names as a hosts file, node-zones as a resolver's zones. Every
new consumer (ssh's known_hosts, an authorized_keys) meant another formatter
in the control plane, in the consumer's own configuration language.

Now a fact is a path and a Go template over the roster view (this node, the
suffix, and every served name vs the machines). The mesh owns the data; the
module owns the format. /etc/hosts is a template on the network module;
dnsmasq's zones move to dnsmasq. The controller renders and reads neither.

WireGuard stays a computed generator: the overlay is the substrate delivery
rides on, and its config is topology, not a roster projection.

Output is byte-for-byte unchanged, pinned by the hosts golden tests and the
resolver tests that compose the real dnsmasq manifest.
2026-09-27 01:33:20 +02:00
jschoubben 7180a273a2 The enrolment user is per token, and it has an inbox
Design 25 §6 says an enrolling node subscribes the inbox its own token derives.
It had none: `sub` was empty, so a node would publish its request and wait out
its timeout against a mesh that had answered — the handshake could not have
completed.

And there was one shared `enrolment` user, which cannot carry that inbox at all:
a permission belongs to a user, so an inbox per token means a user per token.
Named after the node, which **is** the token's id — a token is issued for a node
record, the mesh holds one live claim per record, and the node's name is the one
identifier both sides have before anything else is agreed. It is also exactly
what the other transport does, where the account is named after the node and the
secret is its password.

A nameless enrolment user is now refused rather than composed into
`_INBOX.enrol..>`: an empty subject token, and worse, one every nameless
enrolment user would share — which is one machine able to read the credentials
sealed to another.

Still to wire: something that composes one of these per live token. Nothing
composes enrolment users yet, on either bus — on the old one the account is made
imperatively through the broker's management API when a token is issued, and
here there is no management API, so issuing a token has to recompose the server's
configuration. That is the remaining half of enrolment on the new bus.
2026-09-27 01:31:34 +02:00
jschoubben e65b3950cc A node's declaration consumer, which only the mesh can make
A host's account reaches no part of the JetStream API — correctly, because the
controller is the only writer of consumer definitions — so the object a node
reads its declarations through has to be waiting before the host binds to it,
and nothing created one. Named after the node, because the node's own ack grant
is `$JS.ACK.NODES.<node>.>` and a consumer named anything else is one the host
cannot acknowledge a delivery from.

No max-deliver, and a five-minute ack wait: a declaration is settled only after
the node has applied it and reported, which is minutes on a machine pulling
images, and the stream holds exactly one message per node — so there is nothing
to dead-letter, only one message to redeliver for as long as that node is away.

Asserted on start as well as created at enrolment, for the reason the streams
are: a mesh raised from a restored backup has node records and no consumers, and
a node whose consumer is missing hears nothing while everything else about it
looks correct.

The test sets the consumer against the grant the node actually gets, because
each of the three ways of getting it wrong is silent: a wrong name cannot ack, a
wrong filter reads another node's declarations, and a pull consumer is one a host
has no authority to bind.
2026-09-27 01:24:46 +02:00
jschoubben a0e09695e5 Merge pull request 'The uplink seat (ADR 0117), and the mesh's names written into the hosts file, not over it (hq 128)' (#79) from feat/the-uplink-seat-and-the-hosts-region into main 2026-09-26 23:00:19 +00:00
jschoubben 88bef39952 The consume side on NATS, and the window held by the server
The other implementation behind the seam, so the store-window guarantee now has
both: one loop, one message at a time, the same window deciding. What differs is
where a held message lives, and that is the whole point of the move — the AMQP
side keeps an unacknowledged delivery in this process, bounded by the prefetch
and lost if the controller stops; this keeps eight bytes saying when the window
opened, and the message stays the server's.

Checked against a running server, seven claims that reasoning cannot answer: a
report is heard and leaves the work queue; one the store cannot take is naked
with a delay, stays in the stream, and is recorded when the store returns; one
about a superseded declaration is settled without being acted on; one the store
never takes is let go once the bound passes; a heartbeat is heard and nothing is
persisted; and the enrolment answer reaches the address the request carried in
its payload — the test design 25 §2 asks for, so the reason for that field
cannot quietly become folklore.

Three things the wiring forced into the open:

**The controller could not have consumed a module event.** Its permissions
granted no event subject to subscribe and no ack subject on the events stream,
so every announcement would have been redelivered for ever, refused by the list
it already had. Both narrow: each followed subject named, not `mesh.mod.*.>`.

**The controller's consumers are not derived.** It files no manifest, so its
authority cannot come from a declaration that does not exist; they sit beside the
mesh's own streams and are asserted the same way. No max-deliver on CONTROL —
the window's bound is the controller's, and a server that dead-lettered first
would discard the push the stream exists to protect.

**Channels, not callbacks.** The library would run a handler on its own
goroutine, and the window's bookkeeping is unlocked because the AMQP loop never
had two.
2026-09-27 00:53:33 +02:00
jschoubben 06cf3c04e5 The consume side behind a seam, and the window wiring into the loop
The outbound half went behind `Bus` and the transport stopped reaching its
callers; this is the other half, and the larger one. Every handler took
`amqp.Delivery`, so the serving loop could not move to another bus without
moving enrolment, reports, builds, upgrades and catch-up with it in one breath.

`Control` states one message in the mesh's words — took it, dropped it, or held
it for the store — and `Inbound` is where messages come from. The AMQP
implementation is today's loop moved rather than changed: same queues, same
prefetch, same holding, because the mesh is running on it and a bus nothing
speaks yet is no reason to alter the one every node is on.

The window (window.go) is now what decides, instead of the conditions that were
inlined in the loop. Two things that surfaced in the wiring:

**Supersession is asked before the store, not after.** A report about a
declaration the mesh has moved past would otherwise wait out a restarting store
to be written and then overwrite what the node is doing now.

**Half of a report is not about a declaration, and that half is never stale.**
What the machine *is* — the tunnel it took over, the ports its own bundle
holds, what an adopted node found, a node moving its overlay key — reaches the
mesh on a report and nowhere else. A rekey set aside as stale is a node whose
overlay key never moves, and no retry is coming, because the node said it once.
So staleness is asked only of a report that is purely an apply's account.

The one thing holding-in-memory can do that holding-in-the-server cannot is
named rather than hidden: `About` sets aside a held message when a newer one
about the same thing arrives, and the bus being built ignores it because the
digest answers the same question.
2026-09-27 00:44:16 +02:00
jschoubben 7a8a19b11b A person's account (step 4.4, the account half)
Design 25 §7. A person is not a module and holds no seat: nothing is
addressed to them, nothing is delivered to them, and they have no durable
consumer. What they have is permission to ask, as a list of tools or `*`
for an administrator.

Four properties the tests hold it to, each of which is a way of being
wrong that would not announce itself: a person reaches nothing but tools,
so one cannot claim a module said something; no ack subject, because
authority over a consumer that does not exist is authority nobody would
audit; no allow_responses, because a person who can answer a request is
impersonating a module on a bus where anyone may serve a tool; and two
people do not share an inbox.
2026-09-27 00:17:52 +02:00
jschoubben ce6ac057f6 gofmt the probe 2026-09-27 00:14:47 +02:00
jschoubben abce68fdf0 Verify that a reply address does not survive a stream
Design 25 §2 builds enrolment around carrying the reply subject in the
payload, because a JetStream consumer claims the transport Reply field for
its own ack address. The whole handshake rests on it, so it is checked:
the caller asked for _INBOX.LCr3M83q... and the consumer saw
$JS.ACK.PROBE.probe_consumer... The design was right, and the workaround
is necessary rather than defensive.

Worth having as a test rather than a note: if a future server version
stopped doing this, enrolment would keep working and the reason for the
payload field would quietly become folklore.
2026-09-27 00:14:32 +02:00
jschoubben cc019908c2 Asking a tool goes through the seam, and loses two problems
Step 3.4. On the new bus there is no reply queue to declare and no
correlation to check: each account is granted one inbox prefix and no
other, so an answer cannot reach the wrong asker. That settles a cost
build.go records having paid — on a shared reply exchange every asker saw
every result, which is why the correlation was checked rather than assumed.

And a tool nobody serves says so at once rather than after the whole wait.
The difference between "that module is down" and "that tool is slow" is the
first thing a person asking wants, and both tests are against a real server
because both are claims about what the server does, not about this code.

RequestBuild stays as it is, and is a different shape on the new bus rather
than the same one: a build takes minutes, so it is work submitted to a
queue with the outcome returning to a reply subject the request carries —
the pattern design 25 §2 already sets for anything crossing a stream. It
touches the builder too, so it goes with that conversion.
2026-09-27 00:11:30 +02:00
jschoubben d0a9abcb1c The store window as a decision, and the problem moving it to the server
introduces

Step 3.4, the consume side's hard part. The guarantee (ADR 0083) is that a
push the controller cannot record because its store is restarting is held
and retried — never dropped, never falsely acknowledged. Keeping the
delivery unacknowledged in memory becomes a nak with a delay: the server
holds it, the controller keeps no list of parked messages, and a controller
that restarts mid-window loses nothing it was holding.

That is a plain win and it introduces one problem. Holding in memory let
the controller drop an older report when a newer one for the same node
arrived, "because acting on it after the newer would undo the newer". A
naked message is the server's and comes back whatever happened meanwhile,
so the older report is redelivered after the newer was applied.

The answer was already in the message. A report carries `Declared`, the
digest of the declaration it is about, which exists because an earlier
attempt to order reports by time lost the race it invited. So supersession
stops being something the controller remembers and becomes something it
checks — the same shape as a node refusing a superseded declaration by
sequence (issue 107): ordering settled by what a message says, not by when
it arrived.

Pure, so the guarantee is testable without a bus, a store or a clock. Nine
tests, including that staleness is decided before the store is waited on —
a redelivery that lost its race must not hold a slot a current message
needs.
2026-09-27 00:05:52 +02:00
jochen 63ba2d178f review: hold the hosts region through composition, and say so in plan
A declaration-level test composes the shipped networking module with a
resolver and asserts /etc/hosts arrives as mesh-wireguard.fact-node-names
with into: block and region-only content, that the resolver's restart-on
still names it, and that a resource's at passes through untouched — a
composition step dropping into would otherwise go unnoticed. plan --show
marks files written into, so a region is not read as the whole file.
The rollout order is spelled out: every node's host, the controller's
own included, must be block-aware before this controller ships (hq 128).
2026-09-27 00:05:41 +02:00
jschoubben 6c12780abe Describe the bus on its own terms
Comments framed the new bus by what it replaces — a comparison in almost
every explanation, which reads as though NATS were a variant of the old
thing rather than the mesh's nervous system. Removed throughout, and
OverAMQP becomes OverCurrent: the seam's two sides are the bus the mesh
runs on today and the one being built, not two protocols.

What remains is the client library's own package name, which is its name.
2026-09-26 23:51:00 +02:00
jschoubben 92d87b0082 gofmt bus.go — import grouping 2026-09-26 23:47:26 +02:00
jschoubben 2fad32767e The controller's outbound link behind a seam, with both transports
Step 3.4, first half. Every one of these took an *amqp.Channel, so the
transport reached every caller and swapping it meant touching all of them.
The seam turned out to be small — the controller sends exactly two kinds of
message that expect no answer — which is the same measurement that said
this bus could be replaced at all.

Bus is stated in the mesh's words, not a transport's: PublishEvent and
PublishDeclaration. Two implementations, both shipping, because steps 1 to
4 leave every node on AMQP and the NATS one is selected at the rollout.
Both ship is also what makes them comparable: one conformance fixture holds
both to the same envelope, and the NATS one is checked against a real
server reading back from the stream rather than from the code that wrote it.

Still on *amqp.Channel: RequestBuild and Ask, which carry reply-queue
machinery, and the whole consume side — the control loop, enrolment, serve.
2026-09-26 23:47:15 +02:00
jochen 51163b9f14 The mesh's names are written into the hosts file, not over it (hq 128)
/etc/hosts is the machine's: the distribution's localhost lines, the
operator's own entries, and marked blocks other tools maintain there.
Writing node-names whole replaced all of it the moment the private
network was taken, and every later write by those tools was lost at
the next machine joining. The node-names fact is now emitted with
into: "block", so the host owns only its marked region and keeps the
rest byte for byte. The region holds only the mesh's names: no header
claiming the file, no localhost, no 127.0.1.1 line — the floor was
never the mesh's to write. How a fact is written is a property of the
fact in the closed table; node-zones stays a whole file the mesh owns.

Sequencing: a host older than the block mode refuses the whole
declaration on an unknown into, so every host must be upgraded before
this controller is rolled out.
2026-09-26 23:46:51 +02:00
jochen 6557aca750 A machine's uplink is a seat (hq ADR 0117)
the-uplink joins the closed set as a node seat delivering nothing. Its
holder is the module for the machine's own network manager, and keeps
that manager from contradicting the mesh — the resolver file left to
resolv-conf, mesh0 left alone — without ever declaring a link. Held
per machine, so a machine running two managers is refused at
assignment rather than found by its resolver being rewritten. The
count test moves to fifteen; one test holds the seat's shape.
2026-09-26 23:45:47 +02:00
jschoubben 1801f1178e Hold the Go emitter to the shared fixtures
Every required header set, each value in the pinned shape, and the subject
derived the same way. Read from the sdk's conformance directory by sibling
path, never copied.
2026-09-26 23:40:59 +02:00
jschoubben f21a84c510 Merge pull request 'The controller marks a deliberately-empty declaration owns_nothing (hq 127)' (#78) from feat/controller-marks-owns-nothing into main 2026-09-26 21:40:18 +00:00
jschoubben e14b02991e The controller marks a deliberately-empty declaration owns_nothing (hq 127)
The host refuses an empty body unless told the emptiness is meant
(mesh-host#29). When a node's declaration composes to no resources —
which #77 now sends rather than skips — Body() sets owns_nothing, so
the node applies it and drops what it last held. A declaration with
resources never carries the marker. One test.
2026-09-26 23:40:04 +02:00
jschoubben bf82805cfd A manifest holds no subject, checked across all 72
Design 29 §1's load-bearing rule: a module names its events, tools and
seats locally and the mesh derives the subject, so reorganising the subject
space leaves every manifest correct. It held by construction, and a rule
held by construction is one a later field breaks quietly.
2026-09-26 23:34:42 +02:00
jschoubben 08ebb8c999 Merge pull request 'An empty declaration is sent, so a node drops what it last held (hq 127)' (#77) from fix/127-an-empty-declaration-is-sent into main 2026-09-26 21:32:43 +00:00
jschoubben 0a8a592ef4 An empty declaration is sent, so a node drops what it last held (hq 127)
push skipped any node whose declaration composed to zero resources. A
node that HELD something before — the broker opening a placement gave
an adopted node, say — then kept it forever: the empty declaration that
would drop it was never sent, and the node's own heartbeat re-applied
the stale resource with no way for the mesh to say it is gone. Now the
empty declaration is sent; the host drops what the mesh owned and keeps
what it found. A node that never held anything applies it as a no-op.
Surfaced on ace: the foundation-opening fix (#74) removed its only
resource, and the correction could not reach it until this.
2026-09-26 23:32:30 +02:00
jschoubben 173c8c7c21 Rename the mesh's seats to mesh-*, keeping their interfaces
novox/hq ADR 0118: the prefix is the reservation rule, so a module
declaring any mesh-* name is refused and there is no reserved-names list to
drift. Ten seats renamed in the table, the manifests that claim them, the
controller's own shipped manifests, and the tests.

Not the migration 0118 expected: a holding is derived at resolution from
manifests and never stored, so nothing recorded points at an old name. A
kept rename table tells a manifest written against one what it became —
kept rather than retired, because a module lives in its own repository and
may be registered long after the catalogue stopped using it.

**A seat is not the interface it delivers.** The git seat became mesh-git
and the git provision did not; likewise the package registry. A blanket
replace renamed both, and the failure read "the package registry is served
on <nil>", which does not say "you renamed an interface". A test now pins
every seat against what it delivers, and that neither name is also the
other.
2026-09-26 23:07:09 +02:00
jschoubben 98d348d5b9 Merge pull request 'The mesh's interface takes over the found tunnel's MTU' (#76) from feat/controller-carries-tunnel-mtu into main 2026-09-26 20:41:00 +00:00
jschoubben 7fc5fd02fd The mesh's interface takes over the found tunnel's MTU
Carries MTU from the reported tunnel (mesh-host#28) through inventory,
the overlay graph's TakeOver, into the generated config's [Interface].
A tuned path keeps its MTU across the takeover instead of regressing to
1420 and hanging transfers no ping would reveal. Two emit tests; a
tunnel with no MTU writes no line.
2026-09-26 22:40:42 +02:00
jschoubben 50878a9d98 Merge pull request 'A taken tunnel brings its ListenPort, even on a node the hub cannot dial' (#75) from fix/a-taken-tunnel-brings-its-port into main 2026-09-26 20:34:10 +00:00
jschoubben cc252472e2 A taken tunnel brings its ListenPort, even on a node the hub cannot dial
A home node behind NAT (no Endpoint → not Reachable) that took over a
tunnel must still listen on that tunnel's port: its LAN peers dial it
there. ListenPort was gated on Reachable, which conflated 'a peer dials
me here' with 'the hub can dial me' — so the takeover guard refused
overlay-up, and the guard's suggested remedy (re-place with an
endpoint) breaks a NAT'd node's path: it stops keepalive and hands the
hub a private LAN address to dial. TakeOver now carries the found
tunnel's port (already known to the controller), and the interface
listens on it when the node is not otherwise reachable. Two tests;
Endpoint-reachable nodes keep the old path unchanged.
2026-09-26 22:33:27 +02:00
jschoubben aa74bd86ca Derive a seat's stream and a module's consumer, and wire JetStream
Task 3.9's other half and 1.4's missing client. The derivation is pure and
unit-tested; only "does the server accept this" needs one running, behind
MESH_TEST_NATS so the ordinary suite stays offline.

A seat's work queue is created at registration, not assignment, so work
queues until a holder appears — a stream created at assignment would make
"the holder is not here yet" mean "your messages are gone". Named after the
seat, because the holder can change and the queued work must not care.

A holder's worker uses a queue group even though the seat guarantees one
holder: the seat is authority, the queue group is delivery, and tying them
together means the day somebody allows two holders every message is
processed twice with nothing reporting it.

One consumer per module carrying every filter, because its ack permission is
derived from its name.

And a real bug the live server caught: a durable name may not contain a dot,
but an ack subject is $JS.ACK.<stream>.<consumer>, so the single string that
read correctly inside the permission was rejected as a consumer name. Split
in two, beside the permission that has to match. Unfixed, the symptom would
have been every message redelivered forever with a permission list that
looks right — which is the failure design 25 §4 warns about.
2026-09-26 22:28:42 +02:00
jschoubben d2ab0b2b82 Merge pull request 'The broker opening is only on the broker's host, not every node' (#74) from fix/foundation-opening-only-on-the-broker-host into main 2026-09-26 20:20:50 +00:00
jschoubben 48d8c89749 The broker opening is only on the broker's host, not every node
Enrolling ace applied adoption.opening-tcp-5671-incoming to it, opening
5671 from anywhere (v4+v6) where nothing listens — the ace session
caught it. foundation ports widen the broker's from:mesh port to
from-anywhere so a machine that is not yet on the mesh can make its
first dial; that belongs on the broker's host alone. foundationPortsFor
keeps the port only when a module resolved onto this node listens on
it, so novox opens 5671 and a node that merely dials out opens nothing.
Two tests, both directions.
2026-09-26 22:20:10 +02:00
jschoubben 7232d6df4b A module may declare its own seats (novox/hq ADR 0118)
The manifest carries seats and uses; registration refuses a mesh-* name, a
duplicate declarer, an undeclared uses or claim, a seat with no protocol, a
scope mismatch, and a holder that does not answer what its seat promises.

The parser stops judging unknown claim names, because it cannot: another
module may declare that seat, and one manifest cannot tell. The test that
encoded the old rule is rewritten to assert the refusal at registration, and
a new one pins the case the parser could not have distinguished.

Tools are declared for the first time, under their own key — serves already
means a provision's facts.
2026-09-26 22:17:17 +02:00
jschoubben 112cbd294d Per-subject caps on EVENTS, and a comment corrected against the server
NATS refuses overlapping streams rather than double-storing, which is the
opposite of what the Overlaps comment claimed. The check still earns its
place — it names both streams at composition rather than one at apply — and
the refusal is what rules out a shared stream beside per-module ones.
2026-09-26 21:49:55 +02:00
jschoubben 553814b6eb The mesh bus seat is one per mesh, and the amqp broker does not contend
Step 2.3 of novox/hq ADR 0116. The refusal is the resolver's existing one;
these pin it for this seat, including that a different bus implementation is
refused for the same reason — the property that lets the bus be replaced.
2026-09-26 21:40:03 +02:00
jschoubben eeb0560fc2 The mesh-broker seat delivers mesh-bus (novox/hq ADR 0120) 2026-09-26 21:17:21 +02:00
jschoubben 1f36787d75 The mesh's four streams, asserted on every start
Task 1.4. The foundation set only — a seat's streams come at registration
and a module's consumers at assignment, neither of which has happened at
genesis (ADR 0118).

Asserted rather than created: a stream that was deleted, or a mesh raised
from a backup, must converge rather than run without the guarantee its
messages assume.

Two things the definitions have to get right, both tested:
- CONTROL names its subjects instead of taking mesh.control.>, because
  heartbeats live under that prefix and a stream of them competes for
  retention with the messages that matter
- EVENTS filters on the event token, which is why that token exists; a
  filter over a module's whole namespace would persist every tool call

Overlapping filters are refused where the set is written: NATS accepts two
streams matching one subject and stores the message twice under two
retentions, which nothing reports.

Adds nats.go as a dependency; it pulled golang.org/x/* forward. Full suite
green.
2026-09-26 21:02:18 +02:00
jschoubben c753f9d5c0 Compose the bus's accounts instead of calling a management API
Task 1.3 of novox/hq ADR 0116. On AMQP an account was an HTTP call; on NATS
it is text the controller composes and the server reloads (ADR 0106). Pure,
so the mesh's whole authority model is testable as strings.

NATS closes a gap management.go recorded rather than hid: LavinMQ has no
topic permissions, so an emitter was granted the events exchange whole and
ADR 0042's origin reservation was "stamped by the sdk, not enforced here".
Per-subject permissions make it the server's refusal.

Two things found by composing a real file rather than reading the design:

- a scoped inbox leaves a responder unable to reply, because the answer goes
  to the caller's inbox. allow_responses is the answer — one reply to the
  subject of a message actually received — and only principals that serve
  are granted it. Recorded in design 25 §4.
- composition must be deterministic: the module's entrypoint reloads on the
  file's digest, so an order-dependent composer would reload the whole bus
  on every controller restart. Covered by a test.

The golden fixture is the exact text `nats-server -t` accepts, so the syntax
is the server's rather than one we invented.
2026-09-26 20:58:49 +02:00
jschoubben 2f7b407000 Merge pull request 'A carried peer is nameable, and the mesh answers for it (hq 112)' (#73) from feat/112-a-carried-peer-is-nameable into main 2026-09-26 18:10:00 +00:00