Commit Graph
485 Commits
Author SHA1 Message Date
jschoubben 7efcccd013 Merge pull request 'The build machine takes work on the bus its credential names, and the work queue has a taker' (#107) from feat/the-build-machine-takes-work-on-nats into main 2026-09-27 23:59:56 +00:00
jschoubben 964285f08c The build machine takes work on the bus its credential names, and the work queue has a taker
Two halves of one gap the first build over the new bus met. The machine decided
its bus from a variable its container never received, so the credential the mesh
sealed to it went unread; a credential for the new bus names the bus by scheme and
carries user, password and fingerprint beside the address, and that is enough to
dial it, pinned. And the roles' work queues were raised with no holders, so the
consumer a machine binds to take work was never created: the holders are read
from the catalogue and the handover record, as the resolver reads them.
2026-09-28 01:59:20 +02:00
jschoubben 5698dda11f Merge pull request 'A principal may hear what its consumer delivers' (#106) from fix/a-principal-may-hear-its-consumer into main 2026-09-27 23:47:29 +00:00
jschoubben 6005a8471f A principal may hear what its consumer delivers
A push consumer delivers on _DELIVER.<its name>, and a client bound to it
subscribes exactly that. No principal was granted it, and the server refused
every one the first time it bound a consumer: the control plane, each machine,
and a module would have been next. Each kind is granted its own consumers'
delivery subjects and no other's. The line announcing the raised bus printed the
URL with the credential in it; the address alone now.
2026-09-28 01:46:16 +02:00
jschoubben e2ee0dfe98 Merge pull request 'The bus account has JetStream, and the control plane's client has its own inbox' (#105) from fix/the-bus-account-has-jetstream into main 2026-09-27 23:40:38 +00:00
jschoubben 70341cfbc7 The bus account has JetStream, and the control plane's client has its own inbox
Two refusals the first live connections met. A user in the MESH account was told
"JetStream not enabled for account" the first time it bound a consumer: with
accounts defined, JetStream is enabled per account, not only globally — the
account's setting, which the mesh owns, not the server's block, which it does not.
And the control plane's client used a random inbox prefix where it is granted
exactly _INBOX.<its user>.>, so the server's first answer could not reach it. The
prefix now follows from the user in the URL, for every principal that dials so.
2026-09-28 01:40:10 +02:00
jschoubben 77643aa3f4 Merge pull request 'The control plane pins the bus's certificate, and keeps its password out of errors' (#104) from fix/the-controller-pins-the-bus-certificate into main 2026-09-27 23:35:56 +00:00
jschoubben 1fd6194ff8 The control plane pins the bus's certificate, and keeps its password out of errors
The bus presents the mesh's own certificate, which names nothing a public verifier
accepts; the client verified by name and failed against a bus that was answering
("certificate is not valid for any names", 2026-09-28). It now pins the leaf's
fingerprint from MESH_BROKER_CERTIFICATE, as every host does. And a connection
error named the whole URL, password included — the address alone now.
2026-09-28 01:35:20 +02:00
jschoubben c37018fdd2 Merge pull request 'The control plane serves and pushes on the bus it is told to' (#103) from feat/the-controller-serves-on-nats into main 2026-09-27 23:30:52 +00:00
jschoubben 3907ea0db0 The control plane serves and pushes on the bus it is told to
The seams were there and nothing chose a side: serve, push, ask and build all
opened the old bus's connection and declared over its channel, whatever
MESH_BUS_NATS said. So the switch moved every host and left the control plane
unable to follow — "this control plane has no MESH_BROKER_AMQP" with the new bus
named and standing (2026-09-28). That was task 4.3 of design 28, still open.

One place now decides: connectLink reads the switch, refuses both buses named at
once, raises the new bus's streams and this controller's consumers when it is
handed the inventory, and opens the link over whichever bus it is on. Every
caller that sent a declaration or asked a tool through the old channel goes
through the server's bus instead, which the new transport has and the channel is
not. OverNats is that outbound: a declaration is a JetStream publish into the
node's own subject, an event is announced on the subject its name derives to, a
tool is request and reply on the module's tool subject.
2026-09-28 01:29:44 +02:00
jschoubben 40f5e9a41c Merge pull request 'A machine may bind its consumer' (#102) from fix/a-node-may-bind-its-consumer into main 2026-09-27 23:18:20 +00:00
jschoubben 2c2eb51878 Only CONSUMER.INFO was missing from a machine's grants; the rest was already there 2026-09-28 01:17:40 +02:00
jschoubben aa2d0b51ea Golden: a machine's user may bind its consumer, ack, and hear its inbox 2026-09-28 01:17:15 +02:00
jschoubben 64d154d9d7 A machine may bind its consumer and hear the answer
Binding to a consumer asks the server about it and hears the answer on the
client's inbox; hearing a declaration acknowledges it. A machine's user was granted
none of that and was refused the first time one dialled a permissioned server:
"this node cannot read its declarations". Its inbox is its own prefix, which the
host now sets.
2026-09-28 01:16:46 +02:00
jschoubben ffa390f916 Merge pull request 'The mint leaves the control plane's old-bus secret alone' (#101) from fix/mint-leaves-the-control-planes-old-secret-alone into main 2026-09-27 23:08:35 +00:00
jschoubben 4d62e6caf1 The mint leaves the control plane's old-bus secret alone
The control plane is a module too, and its broker secret is the old bus's
credential it is still using while the mint runs. Writing the new bus's blob there
cut the mesh off from its own old bus mid-move. Its new-bus credential is the
controller principal's bus secret; the module principal is skipped.
2026-09-28 01:08:11 +02:00
jschoubben 386ae676ca Merge pull request 'The control plane mounts the bus secret it reads' (#100) from fix/the-controller-mounts-its-bus-secret into main 2026-09-27 23:03:42 +00:00
jschoubben f8a9c3d6bc The control plane mounts the bus secret it reads
MESH_BUS_NATS_FILE named /run/secrets/bus and nothing put a file there: the
manifest binds each secret explicitly, and the switch added the secret and the
variable but not the bind. Found live — the control plane came up on the new bus
and could not read its own credential.
2026-09-28 01:03:18 +02:00
jschoubben 83671fae5f Merge pull request 'The network map resolves each machine with the seat holders on record' (#99) from fix/the-network-map-knows-the-holders into main 2026-09-27 22:56:55 +00:00
jschoubben 9b715524a2 The network map resolves each machine with the seat holders on record
Without them, a machine running the next holder of a seat beside the current one
resolves as two holders, is refused, and drops out of the map — and with it the
address every other machine composes for what it offers. Found live: the control
node vanished from the private network the moment the new bus was assigned beside
the old one, and nothing on any machine could be composed.
2026-09-28 00:55:39 +02:00
jschoubben e06fc1ed16 Merge pull request 'The control plane speaks the new bus' (#98) from switch/the-controller-speaks-nats into main 2026-09-27 22:38:27 +00:00
jschoubben 13d7c5c5dd Merge pull request 'The mint names the bus by bare host, and can mint again' (#97) from fix/mint-bare-host-and-again into main 2026-09-27 22:37:36 +00:00
jschoubben 7b02feaebb The mint names the bus by bare host, and can mint again
BareAddress adds a scheme where none was, so the host it yielded carried one and
every URL built from it carried two. Caught before a push: the host is now taken
with no scheme and no port, and every URL adds its own. `rollout mint --again`
mints every credential afresh for exactly this case — a mint that was wrong before
anything received it.
2026-09-28 00:37:10 +02:00
jschoubben bd10e2c695 Merge pull request 'The mint finds the bus on the hub' (#96) from fix/mint-finds-the-bus-on-the-hub into main 2026-09-27 22:34:41 +00:00
jschoubben 4b209d944d The mint finds the bus on the hub
The map of who is where on the private network lists the machines placed around
the hub, not the hub — and the control node is the hub, and runs the new bus.
Found on the first live run: refused for having no address. The address every
machine already dials the current bus at is the same machine, so that host with
the new port is what they are told.
2026-09-28 00:34:17 +02:00
jschoubben 84024cbdb6 The control plane speaks the new bus
One environment variable moves it (design 25): MESH_BUS_NATS, read from the `bus`
secret `rollout mint` sealed to its machine, and the two that named the old bus go,
because being told about both is refused at start. Registered and built ahead of
the push that flips it, so the flip and the bus it flips to arrive in one
declaration — the same push that starts the new bus and hands this machine its
membership. Deliberately not merged until every credential is minted.
2026-09-28 00:31:57 +02:00
jschoubben ad2eed2f71 Merge pull request 'The move mints every credential and tells each machine its membership' (#95) from feat/the-move-mints-and-delivers into main 2026-09-27 22:30:58 +00:00
jschoubben e8aa7ed9e7 The move mints every credential and tells each machine its membership
`rollout mint` gives every principal the new bus will have a credential it does
not yet have and puts each where its owner reads it: a machine's as a membership
— bus address, fingerprint, password, transport — sealed into its declaration
(migration 0041, the `bus-membership` resource the host reads after applying); a
module's as its broker secret, through the same delivery `module issue` uses; the
control plane's own as its `bus` secret. Idempotent, and worked out from where the
bus's module is assigned rather than from this process's environment, because this
process is still on the old bus when it runs and must be.

This is the half of design 28 task 5.2 the first live attempt found missing: a
credential was minted only at enrolment, at `module issue` and for a person, so no
machine already enrolled could ever be moved. `rollout check` was right to refuse;
now there is something to run first.
2026-09-28 00:16:24 +02:00
jschoubben 337aaea123 Merge pull request 'The first handover records the standing holder without re-judging it' (#93) from fix/record-the-standing-holder into main 2026-09-27 21:35:08 +00:00
jschoubben 4c41628b20 The first handover records the standing holder without re-judging it
On a mesh that predates the record, every handover has to begin by writing down who
already holds the seat — otherwise the next holder cannot be assigned beside it,
because two eligible claimants with nothing on record are refused. Found on the
live mesh minutes after 0040 moved the bus seat's row: the standing holder no
longer satisfies what the seat delivers, on purpose, and so could not be recorded,
and so nothing could stand beside it.

Recording who already holds is not making a new holder. Derivation never read
what the seat delivers, so the standing holder holds regardless; when nothing is
on record and the named assignment claims the seat at its scope, only that claim
is checked. Every change of holder is still judged in full.
2026-09-27 23:34:40 +02:00
jschoubben ae7fb520d7 Merge pull request 'AMQP is not a provision: the bus seat delivers the bus, and the word is refused' (#92) from feat/amqp-is-not-a-provision into main 2026-09-27 21:30:09 +00:00
jschoubben f325073982 AMQP is not a provision: the bus seat delivers the bus, and the word is refused
Two halves of novox/hq ADR 0131. Migration 0040 moves the mesh-broker row from
`amqp` to `mesh-bus`, so the seat's holder answers for the mesh's bus and not for
the wire protocol the old broker spoke — which is what let only the retiring
broker hold the seat that names the bus. Safe under the current holder: the
control plane composes its own address through the seat by name and the overview
derives holders by name; only registration and provision resolution read the
column. What must not happen in between is re-registering the current holder.

And the parser refuses a manifest that provides or requires `amqp`, each refusal
saying what to do instead: a module reaches the mesh's bus through the sdk and
depends on the seat, not on a protocol. A whole-catalogue test asserts nothing
beside this checkout names it; the three modules that did are removed there.

Two tests that used the old broker as a fixture now use the module that replaces
it or a manifest this package owns.
2026-09-27 23:29:00 +02:00
jschoubben 33c4e4be34 Merge pull request 'A seat is handed over as one act, and the holder is on record' (#91) from feat/seat-handover into main 2026-09-27 21:22:50 +00:00
jschoubben 585a6abbdd A seat is handed over as one act, and the holder is on record
`seat <name> --to <node>/<module>` makes one assignment the holder of a seat in
the same write that removes the previous one. The row is new (migration 0039);
without one, the resolver derives the holder as it always did — the sole eligible
assignment, two refused — so nothing changes for a mesh that never hands a seat
over. With one, the recorded assignment holds and any other whose module could
hold the seat is eligible and silent: not refused, not holding. That is what lets
the next holder run beside the current one until the switch (hq design 26, design
28 task 5.3, ADR 0131).

Why: the controller finds its own bus through a seat, and the day that seat was
left with nobody in it — because two eligible holders could not coexist and the
old one's claim was taken away — the control plane looped for two hours while
every service stayed up. A handover that is never empty in between is the fix,
not a workaround for it.

`CanHold` is the one judgement of whether a module may hold a seat — claims it at
its scope, provides what it delivers, against the store's row — shared by
registration and the handover so they cannot drift apart. The holding belongs to
the assignment and goes when it does, so a seat never points at nothing running.

Tests: the resolver with and without a record, on the same and another machine,
under a former name; the store's row replaced not added, refused for an
unassigned target, removed with its assignment; CanHold's four answers and that
they follow the store. Full suite green against a real NATS and store.
2026-09-27 23:22:20 +02:00
jschoubben 8d52a2cfb0 Merge pull request 'The move ends with the old broker going, not staying' (#90) from fix/rollout-retires-the-old-broker into main 2026-09-27 21:05:47 +00:00
jschoubben 1cfe6be9c4 The move ends with the old broker going, not staying
`rollout check` said the old broker stays running as an ordinary provider of
amqp, and this was not its retirement. That was ADR 0127, which ADR 0131 has
superseded: AMQP is not a provision, so once every machine reports on the new bus
nothing of the mesh speaks to the old broker and its module is unassigned. The
plan says so, as its last step.

The flag that made the "it stays" line conditional is gone with the line — there
is no case in which the broker is kept. The test that pinned the opposite now
pins this, and says which record changed under it. The stale citation of a
record numbered 0119 is corrected while here.
2026-09-27 23:04:59 +02:00
jschoubben 4e4481b6f2 Merge pull request 'The store owns the seat set, so only the control plane may judge a claim' (#89) from fix/the-store-owns-the-seat-set into main 2026-09-27 20:00:05 +00:00
jschoubben 63ca073938 The store owns the seat set, so only the control plane may judge a claim
A claim on a seat was checked against `SeatNamed` inside `ParseManifest`, and the
build machine parses manifests too. It has no store, so there it answered from the
set compiled into the binary — a copy of data the control plane owns (ADR 0122).

When the two disagreed, that copy won where it mattered. The store's row said the
bus seat answers for `amqp`; the binary's said `mesh-bus`; and a holder that
provides `amqp` was refused at build time for not providing `mesh-bus`. The seat
went unheld, the controller lost the address it composes through that seat, and the
control plane crash-looped on a bus that was healthy the whole time.

So the two checks that read the set — a seat's scope, and what its holder must
provide — move to CatalogueProblems, which runs only in the control plane and only
after UseSeats has replaced the set with the store's. The parser keeps what it can
judge from the manifest alone, the reserved-namespace rule included.

A test pins it: the same manifest, two different values in the store, and the answer
follows the store both times. It fails if the check moves back.
2026-09-27 21:59:40 +02:00
jschoubben b244a768a3 Merge pull request 'The Go base has to be 1.26: the NATS client requires it' (#88) from fix/go-126-base into main 2026-09-27 19:00:32 +00:00
jschoubben 81d52719f4 The Go base has to be 1.26: the NATS client requires it
`nats.go v1.54.0` declares `go >= 1.26`, so `go mod tidy` puts the directive at
1.26.0 and will not leave it at 1.25. The base image is pinned by digest at
1.25.14-alpine, so nothing in this repository compiles against it — caught when the
build machine's own build failed with "go.mod requires go >= 1.26.0 (running go
1.25.14)".

Moved to 1.26.8-alpine by digest, same flavour as before. The pin stays a digest:
the point of pinning is that the compiler does not change underneath a build, and
that is still true one version along.

Two other manifests carry the same pin for the same reason — they compile this
repository's code — and move with it in the catalogue.
2026-09-27 20:58:08 +02:00
jschoubben f6a93fe74c Merge pull request 'The bus on NATS: both transports behind seams, and the rollout switch' (#87) from feat/nats-genesis into main 2026-09-27 17:36:41 +00:00
jschoubben 497f8ea567 Merge main: the trunk renamed the seats and made them data
Both branches changed the seat set from the same starting point, so every number
collided and every `mesh-*` name existed twice. The trunk's numbers and names win:
this branch's records became 0129/0130 and its migrations 0037/0038, and the
hardcoded rename map gave way to the trunk's `seat_alias` table — a rename is a
row now (ADR 0122), not a recompile.

Three of my checks were wrong and the merge is what showed it:

A seat with an empty protocol is a marker, not an incomplete declaration. Most
node-scoped seats are markers — which module is this machine's packet filter —
and refusing one refused most of the set, the showcase module included. A
mistyped field name is already refused by the parser, so an empty protocol was
written as one deliberately.

A claim on a seat this manifest does not declare is not the parser's to judge. A
module may hold a seat another module declared; that is the whole reason ADR 0126
has callers name the seat and not its provider. Whether the seat exists is a fact
about the catalogue, so the refusal is at registration, where every declaration
is in view.

And a seat may share a name with the provision it delivers. `git`, the npm
registry and the artifact store still do, because renaming a delivering seat
cascades to every consumer requiring it, with a window where a holder stops
resolving mid-flight. The trunk deferred exactly those three on purpose.

Full suite green against a real NATS and store.
2026-09-27 18:50:18 +02:00
jschoubben dded086b54 rollout check: whether this mesh could move its bus, and what is missing
The rollout moves every node at once, so there is nothing to inspect afterwards and no
half to roll back — either the mesh was ready or it was not. That makes the readiness
question the valuable half: it costs nothing, it can be asked of a mesh that is serving as
many times as you like, and every answer is a thing somebody can go and fix.

It reads from records and dials once. Is a bus answering, does a machine hold the seat, has
that machine been sent the composed user list, does every machine have a credential for the
new bus, does every module that speaks. Each missing thing names its own next step, because
"not ready" that cannot be acted on is not an answer — and this is read at the point where
the next step is irreversible.

**A machine with no credential is the one that must stop it.** It keeps running and cannot
come back, and afterwards there is no bus to tell it anything over, so the remedy has to
happen first. The message says so.

A module that never reaches the bus is not counted as missing a credential. A third of the
catalogue never speaks, and listing those would bury the ones that matter.

`rollout --confirm` refuses and says why: the move is not being written before its check has
been run against a real mesh. And the plan it prints says the old broker stays — it remains
an ordinary provider of `amqp` for whatever else uses it, which on this installation is a
whole automation layer that has nothing to do with the mesh. This move is not its retirement,
and that is why it is survivable: what breaks if it goes wrong is the mesh's ability to
change things, not the services its modules serve.
2026-09-27 17:59:01 +02:00
jschoubben e7b5100324 Merge pull request 'The mesh owns the operator's ~/.ssh: account fact + home-scoped resources (to-be 29)' (#86) from feat/to-be-29-mesh-owns-ssh into main 2026-09-27 15:51:55 +00:00
jschoubben 8ceec32692 The mesh owns the operator's ~/.ssh: account fact + home-scoped resources (to-be 29)
A node carries its operator account (name + home; migration 0036, Node.Account,
SetAccount, 'node account' CLI). The account and its home are offered as
machine facts ${machine:account} / ${machine:account-home}, and machineInto
now resolves placeholders in a resource's path and owner (not just content), so
a module writes into a person's home naming what it cannot know. A RosterFile
gains Home: the file is placed under the account's home and chowned to it, its
template sees each node's Account, and a machine with no account gets none —
this is how the ssh Host blocks for every node reach a person's ~/.ssh. Roster
carries per-node accounts (Rendering.Accounts). Tested, including ssh-client
composed end-to-end. Not deployed.
2026-09-27 17:50:56 +02:00
jschoubben 5fcde512bc A build is announced under both names on the old bus, or merging breaks the live mesh
Found by asking what merging this would do to the mesh that is actually running — the
only place the question could have been asked, because the tests were green and both
buses were self-consistent.

Moving the build outcome to the role means a catalogue built from the current manifests
listens for the role's name. The catalogue *already running* listens for the module's,
because that is what it was told when it was installed. The two do not meet, so merging
as it stood would have stopped the live mesh's module graph being updated — silently,
since a binding that matches nothing is not an error.

A rename on a live bus needs the publisher and the subscriber to change together, and a
deployment cannot promise which arrives first. So the old bus announces under both names
and the order stops mattering. The module's own name retires with the bus, in step 5's
list; nothing has ever run on the bus being built, so there is no legacy name there and
this doubling has no counterpart.
2026-09-27 17:39:39 +02:00
jschoubben e5007a7daa The user list is composed before anything moves onto the bus
Found by reading the live mesh's own notes before touching it, which is where this
was heading next.

Composing the bus's user list was gated on the controller already being on the new
bus. That cannot work: the server needs its user list *before* anything moves onto
it. Step 2 of the whole change is exactly that — the server stands in the mesh
carrying nothing, on its own ports, while every node stays where it is. Under the
old gating that step was impossible: the module would come up, find no accounts
file, and its entrypoint would wait for one the controller had decided not to write.

So the only question is whether this machine runs the module that asked for the
file. A mesh that never moves has written a user list nothing reads, costing a few
hundred bytes on one node. The reverse cost a step that could not be taken.

Pinned by a test over the records of a mesh mid-change: everything running, nothing
on the new bus, and a user list that contains the controller — because a file
without it is a bus its own writer cannot connect to.
2026-09-27 17:33:52 +02:00
jschoubben cb77f35a27 A module may watch a role's events, and the catch-up turns out to be unnecessary
Moving the build outcome onto its role broke the one module that consumes it, and my own
agreement check passed anyway. The catalogue's subscription derived
`mesh.mod.mesh-build-machine.event.built` — a module namespace for a role's event, which no
such module owns — so it started, connected, and its graph stayed empty. The check compared
names, and the names agreed: the build machine does emit `built`. Only the subjects
disagreed, and a subscription that matches nothing is silence.

A consumed name is a module's event unless it names a role, and this package cannot tell by
looking — so whoever resolved the declaration says which, the way it already does for a seat
held or used. A module that watches a role gets the role's event subject and a consumer
filtered on it; watching grants subscribe and nothing else, because hearing what a role
announced is not taking part in it.

The check now compares the two halves that actually have to match — the subject a consumer
subscribes against the subject an emitter publishes — with a case pinning that it catches
this exact confusion. Comparing names was checking the easy half.

**And that answered the open question about catch-up: there is nothing to build.** The
mechanism exists because a queue on the old bus receives only what is published after it is
bound, so everything built before the catalogue existed was announced to nobody. A stream is
a log and a consumer is a position in it: a consumer created afterwards starts at the
beginning, so the builds are simply there. Asked of a real server, since the whole decision
rested on it — three builds published with nothing listening, then a consumer created, and
all three waiting for it.
2026-09-27 17:22:30 +02:00
jschoubben 71dbca6392 Merge pull request 'A module declares its fail2ban jail; mesh composes per node (to-be 31 mechanism)' (#85) from feat/a-module-declares-its-jail into main 2026-09-27 15:21:03 +00:00
jschoubben 47d412e13f A module declares its fail2ban jail; the mesh composes them per node (to-be 31)
The mechanism, mirroring Filtering: a module declares Jails (name, failregex,
jail stanza) naming no node/path (ADR 0112); the intrusion-prevention holder
declares Jailing (where composed jails go); the mesh gathers every assigned
module's jails into one jail.d file (a fixed id the fail2ban service restarts
on) plus a filter.d file per jail. A node not running a module has none of its
jails. Tested. Behaviour-neutral until a service module declares a jail — the
per-service content (postgres/mssql/mailu failregex+logpath) is authored next,
against how each container actually logs.
2026-09-27 17:20:50 +02:00