Commit Graph
319 Commits
Author SHA1 Message Date
mesh-admin ad97297576 Merge pull request 'The seat declares the facts its holder states' (#131) from fix/the-seat-declares-the-facts-its-holder-states into main 2026-09-28 14:37:05 +00:00
jschoubben 683b1ed693 The seat declares the facts its holder states
The grant permitted the control plane to state what it applied and the seat said nothing about it, so
the check that every derived subscription has an owner found the catalogue subscribing to a subject
nothing publishes — which is exactly the fault that check exists for, pointed at me.

A seat carries the protocol of its role (novox/hq ADR 0129), so the facts are the mesh-controller
seat's `emits`. That is also what lets another module declare it consumes them. No accepts, so no work
queue is raised for the seat — only what its holder may say. The agreement test now holds all three
places to one another: the seat, the grant, and the words the mesh states them with.
2026-09-28 16:36:45 +02:00
jschoubben 1c3f44a526 Two faults found on review
A second Accept value would have overwritten the first, because the header was Set per value rather
than Added. One value is all any caller passes today, so nothing was wrong — but a helper that
quietly keeps only the last of what it was given is a trap for whoever passes two.

And a replayed announcement that could not be written was published as an empty body: a fact on the
mesh that says nothing, which the reader can only log and drop. It is now said and skipped, because a
body that cannot be marshalled is this program's fault rather than the bus's.
2026-09-28 16:32:41 +02:00
mesh-admin 89e152dfe2 Merge pull request 'The mesh says what it applied, and the replay has an address it may use' (#129) from feat/the-mesh-says-what-it-applied into main 2026-09-28 14:07:20 +00:00
jschoubben 1ebad3786c The mesh says what it applied, and the replay has an address it may use
The pipeline was observable from a merge to an artifact and went dark where it touched a machine: a
node's report is control traffic only the control plane reads, so nothing said which version a
machine runs, or that it refused to (novox/hq ADR 0134). The control plane now states both under the
seat it holds — a role's events belong to the role and keep their address when the holder is
replaced — and only when the report is news, because a machine reconciles every minute and a fact per
report would be a fact per minute per machine.

Whether a report is news is the store's answer: it holds the previous one, so the listener returns it
and the server states the fact. That also gives the catch-up replay a subject the controller may
publish: it was published as a module's event from a module called "control-plane", which does not
exist, so the controller's own account refused it and every catalogue that asked what it missed was
answered with nothing.
2026-09-28 16:07:18 +02:00
jschoubben 4b4c7e0e0d A module's name may contain a dot, so the derived step adds none
A resource's id is `<module>.<its own id>` and a module's name may itself contain a dot — novox.be is
one — so the owner of a resource is everything before the *last* dot. The preparation step's id used a
dot, which made its owner unreadable by that rule; it uses a hyphen, and the id says what it belongs
to whichever way a reader splits it.
2026-09-28 15:44:18 +02:00
jschoubben 338d033632 A manifest HEAD says what it accepts, or the registry answers 404
The check that skips copying a base the mesh already holds asked with no Accept header, and a
registry answers a manifest only in a media type the caller named: the same digest answered 200 with
the manifest types and 404 without them. So the builder concluded it held nothing, copied every
vendor base again, and exhausted the public hub's pull limit a second time today.

The test could not have caught it, because the fake registry answered a manifest HEAD regardless of
Accept — more permissive than the thing it stands in for. It is now as strict as a real registry, and
fails without the fix.
2026-09-28 13:02:08 +02:00
jschoubben 5a963aec10 A version prepares its state before it runs
The mesh derives the preparation from the module's own resource instead of each module hand-writing
a step beside it (novox/hq ADR 0135). A manifest says one word — `prepares` — and the mesh runs that
module's own program in its preparation mode, in the module's own context: the same image, the same
environment, the same mounts, because it is the same code. A published port and a fixed address are
taken away rather than copied, since the version being replaced still holds them.

One word for every kind of module: a Go binary receives `prepare` as its argument, a bundle receives
it through the runtime whose entry takes the same word. The control plane answers it like anything
else — its own schema stops being a special case, and its hand-written step is gone.
2026-09-28 12:43:15 +02:00
jschoubben 2f3bfda8c0 Work slower than the window says so, and one address is the bus's
Three faults the mesh's own logs showed this morning. A handler that outlives the acknowledgement
window was handed its message again while it was still working: acting on a merge builds modules,
minutes against a thirty-second window, so one merge ran the whole catalogue five times over. The
transport now says the work is in progress while it runs, which is where the window belongs.

Everything the mesh hands out — a token, a membership, a person's credential — took its address
from the enrolment setting, which on a mesh that has moved still names the broker it moved from:
the first person issued after the move was handed the retired broker's port. There is one bus, and
its address is the one the control plane is connected to.

And `operator issue` documented an argument order its parser refused.
2026-09-28 09:51:50 +02:00
jschoubben aa771616bb A merge rebuilds what it changed, and what packages it
Three faults in one path. A merge rebuilt every module built from the repository, so one change in
a repository holding twenty-six of them meant twenty-six builds. A merge into a repository a module
only *packages* source from rebuilt nothing — two modules are built from the control plane's own
repository and neither had ever been rebuilt when it moved — because the manifest the mesh keeps
carries no build section, so a build now says which repositories it read and the mesh keeps that
beside what it stood on. And a module handed over by hand could record a repository with no
directory inside it, which is a module nothing can ever rebuild (novox/hq 04-ISSUES/131, /132).

A change inside no module's own directory is a change to what they share, and everything built from
that repository is rebuilt: rebuilding too much is the safe direction, because the fault this whole
path exists for is a mesh that believes it is current and is not.
2026-09-28 09:20:01 +02:00
mesh-admin 1513bbaac9 Merge pull request 'An older merge does not move a source' (#120) from fix/an-older-merge-does-not-move-a-source into main 2026-09-28 03:12:40 +00:00
jschoubben 0014984116 An older merge does not move a source
The forge announces what it finds merged, and an old merge surfacing late moved the recorded head
backwards and rebuilt everything built from that repository, once per old merge. A merge made
before the source was last seen is history; one that says nothing about when is taken as news.
The catalogue now carries when each source was last seen. The bus-records test follows #116:
a module's tools are every one under its own name.
2026-09-28 05:12:37 +02:00
jschoubben 3756bb3460 A base the registry already holds is not pulled from upstream again
A base is named by digest, and a digest the mesh's registry holds under the module's repository
is the same bytes whatever upstream would say. Asked on every build, the public hub's anonymous
pull limit was reached on the first merge that rebuilt a whole catalogue, and every module whose
base lives there failed on a copy it did not need.
2026-09-28 04:48:32 +02:00
jschoubben da31bcb11e A module hears what it consumes: its consumer is raised with the bus, and it pulls it
Every module moved onto the bus by the rollout was issued on the old one, so none had a consumer
waiting; and the grant named a push delivery a runtime's client never binds, while the pull it
does make — asking about its consumer, asking it for messages — was refused. The consumers a
module's declarations imply are now raised whenever the bus is, and the grant is the pull.
2026-09-28 04:29:32 +02:00
jschoubben 0baf727f36 The controller may ask any module's tool
The control plane is the way in for tool calls (novox/hq ADR 0095): a person or an agent asks
through it, so it alone may publish to every module's tool subject. The first ask on the new bus
was refused the publish.
2026-09-28 04:21:25 +02:00
jschoubben 83a298e7e0 A module serves every tool under its own name, and may answer
Every module that served a tool was refused the subscription on the new bus: the grant listed
tools from a manifest field no module fills, because the tools a module serves are what its code
answers and a second copy of that list would be a second source of truth. The grant is now the
module's own tool namespace; nothing else may subscribe it, a caller is still granted per tool by
name, and a module may answer what it was asked.
2026-09-28 04:14:47 +02:00
jschoubben 35252af665 A build records the bases it was handed, and the mesh reads its edges from builds
Bases reach a recipe as build arguments, so the digest was never in the file the builder read
edges from: no build on the mesh recorded what it stood on, and 'build --on', the bases-first
order and the merge follow-up all walked a graph with no edges (novox/hq 04-ISSUES/131). The
builder now reports every base it resolved; the controller records them by artifact path and
reads the newest build's edges from the store, since a recorded manifest carries no build.on.
2026-09-28 03:51:14 +02:00
jschoubben aecac5bda2 One bus: the AMQP transport is gone from the controller
The mesh runs on the seat's bus alone (novox/hq ADR 0131, design 28 task 5.5). The old
transport's consume loop, build request, tool ask, management API and account scoping are
deleted, and the bus switch with them; the controller connects to the broker seat and to
nothing else. The store-window tests keep their assertions on a bus-less fake, and the tests
that only made sense for the old transport's in-memory holding go with it.
2026-09-28 03:36:16 +02:00
jschoubben 81e76fa485 The controller follows the subject it decodes
The decoder named the forge's merge subject as the fourth thing followed and the
list was three long: every message that fell through to that switch panicked the
control plane (2026-09-28). The entry was written and lost between two attempts
at the same edit. A test now walks the list; the composed grants and the genesis
template carry the subject.
2026-09-28 03:13:57 +02:00
jschoubben 525f10b858 A merge on the forge builds what it moved, bases first
The controller follows the forge's merges (novox/hq 04-ISSUES/131). For each
module recorded as built from that repository and branch it records the move to
the merge commit and builds it — bases first, because a module built before the
module it stands on is built against the old one and reports success, and a base
that fails stops what stands on it. Nothing is pushed here: what a finished build
does to the machines running the module stays the upgrade's decision.

Two more things the same ordering gives: `build --behind` builds bases first, and
`build --on <module>` rebuilds everything that stands on a module — the rebuild a
changed base needs, which "behind" does not see because their sources did not
move.
2026-09-28 02:59:02 +02:00
jschoubben c5dc7e732a A store row keeps its seat's protocol, and a holder may take work from its queue
The seat table has name, scope, delivers and decision, and the protocol ADR 0129
gave a seat lives only in the compiled defaults; loading the rows dropped it, so
no role's work queue was ever raised and the first build submitted over the new
bus met "no response from stream". Until the table gains the columns, a row with
no protocol keeps the compiled one of its name. And the holder of a seat is
granted what taking work from its queue needs — asking about the worker consumer
it binds, and acknowledging on it — which the first machine to try was refused.

The control plane's own seat placeholders no longer include the old bus's port,
which the switch removed with the variable.
2026-09-28 02:08:56 +02:00
jschoubben 964285f08c The build machine takes work on the bus its credential names, and the work queue has a taker
Two halves of one gap the first build over the new bus met. The machine decided
its bus from a variable its container never received, so the credential the mesh
sealed to it went unread; a credential for the new bus names the bus by scheme and
carries user, password and fingerprint beside the address, and that is enough to
dial it, pinned. And the roles' work queues were raised with no holders, so the
consumer a machine binds to take work was never created: the holders are read
from the catalogue and the handover record, as the resolver reads them.
2026-09-28 01:59:20 +02:00
jschoubben 6005a8471f A principal may hear what its consumer delivers
A push consumer delivers on _DELIVER.<its name>, and a client bound to it
subscribes exactly that. No principal was granted it, and the server refused
every one the first time it bound a consumer: the control plane, each machine,
and a module would have been next. Each kind is granted its own consumers'
delivery subjects and no other's. The line announcing the raised bus printed the
URL with the credential in it; the address alone now.
2026-09-28 01:46:16 +02:00
jschoubben 70341cfbc7 The bus account has JetStream, and the control plane's client has its own inbox
Two refusals the first live connections met. A user in the MESH account was told
"JetStream not enabled for account" the first time it bound a consumer: with
accounts defined, JetStream is enabled per account, not only globally — the
account's setting, which the mesh owns, not the server's block, which it does not.
And the control plane's client used a random inbox prefix where it is granted
exactly _INBOX.<its user>.>, so the server's first answer could not reach it. The
prefix now follows from the user in the URL, for every principal that dials so.
2026-09-28 01:40:10 +02:00
jschoubben 1fd6194ff8 The control plane pins the bus's certificate, and keeps its password out of errors
The bus presents the mesh's own certificate, which names nothing a public verifier
accepts; the client verified by name and failed against a bus that was answering
("certificate is not valid for any names", 2026-09-28). It now pins the leaf's
fingerprint from MESH_BROKER_CERTIFICATE, as every host does. And a connection
error named the whole URL, password included — the address alone now.
2026-09-28 01:35:20 +02:00
jschoubben 3907ea0db0 The control plane serves and pushes on the bus it is told to
The seams were there and nothing chose a side: serve, push, ask and build all
opened the old bus's connection and declared over its channel, whatever
MESH_BUS_NATS said. So the switch moved every host and left the control plane
unable to follow — "this control plane has no MESH_BROKER_AMQP" with the new bus
named and standing (2026-09-28). That was task 4.3 of design 28, still open.

One place now decides: connectLink reads the switch, refuses both buses named at
once, raises the new bus's streams and this controller's consumers when it is
handed the inventory, and opens the link over whichever bus it is on. Every
caller that sent a declaration or asked a tool through the old channel goes
through the server's bus instead, which the new transport has and the channel is
not. OverNats is that outbound: a declaration is a JetStream publish into the
node's own subject, an event is announced on the subject its name derives to, a
tool is request and reply on the module's tool subject.
2026-09-28 01:29:44 +02:00
jschoubben 2c2eb51878 Only CONSUMER.INFO was missing from a machine's grants; the rest was already there 2026-09-28 01:17:40 +02:00
jschoubben aa2d0b51ea Golden: a machine's user may bind its consumer, ack, and hear its inbox 2026-09-28 01:17:15 +02:00
jschoubben 64d154d9d7 A machine may bind its consumer and hear the answer
Binding to a consumer asks the server about it and hears the answer on the
client's inbox; hearing a declaration acknowledges it. A machine's user was granted
none of that and was refused the first time one dialled a permissioned server:
"this node cannot read its declarations". Its inbox is its own prefix, which the
host now sets.
2026-09-28 01:16:46 +02:00
jschoubben e8aa7ed9e7 The move mints every credential and tells each machine its membership
`rollout mint` gives every principal the new bus will have a credential it does
not yet have and puts each where its owner reads it: a machine's as a membership
— bus address, fingerprint, password, transport — sealed into its declaration
(migration 0041, the `bus-membership` resource the host reads after applying); a
module's as its broker secret, through the same delivery `module issue` uses; the
control plane's own as its `bus` secret. Idempotent, and worked out from where the
bus's module is assigned rather than from this process's environment, because this
process is still on the old bus when it runs and must be.

This is the half of design 28 task 5.2 the first live attempt found missing: a
credential was minted only at enrolment, at `module issue` and for a person, so no
machine already enrolled could ever be moved. `rollout check` was right to refuse;
now there is something to run first.
2026-09-28 00:16:24 +02:00
jschoubben f325073982 AMQP is not a provision: the bus seat delivers the bus, and the word is refused
Two halves of novox/hq ADR 0131. Migration 0040 moves the mesh-broker row from
`amqp` to `mesh-bus`, so the seat's holder answers for the mesh's bus and not for
the wire protocol the old broker spoke — which is what let only the retiring
broker hold the seat that names the bus. Safe under the current holder: the
control plane composes its own address through the seat by name and the overview
derives holders by name; only registration and provision resolution read the
column. What must not happen in between is re-registering the current holder.

And the parser refuses a manifest that provides or requires `amqp`, each refusal
saying what to do instead: a module reaches the mesh's bus through the sdk and
depends on the seat, not on a protocol. A whole-catalogue test asserts nothing
beside this checkout names it; the three modules that did are removed there.

Two tests that used the old broker as a fixture now use the module that replaces
it or a manifest this package owns.
2026-09-27 23:29:00 +02:00
jschoubben 33c4e4be34 Merge pull request 'A seat is handed over as one act, and the holder is on record' (#91) from feat/seat-handover into main 2026-09-27 21:22:50 +00:00
jschoubben 585a6abbdd A seat is handed over as one act, and the holder is on record
`seat <name> --to <node>/<module>` makes one assignment the holder of a seat in
the same write that removes the previous one. The row is new (migration 0039);
without one, the resolver derives the holder as it always did — the sole eligible
assignment, two refused — so nothing changes for a mesh that never hands a seat
over. With one, the recorded assignment holds and any other whose module could
hold the seat is eligible and silent: not refused, not holding. That is what lets
the next holder run beside the current one until the switch (hq design 26, design
28 task 5.3, ADR 0131).

Why: the controller finds its own bus through a seat, and the day that seat was
left with nobody in it — because two eligible holders could not coexist and the
old one's claim was taken away — the control plane looped for two hours while
every service stayed up. A handover that is never empty in between is the fix,
not a workaround for it.

`CanHold` is the one judgement of whether a module may hold a seat — claims it at
its scope, provides what it delivers, against the store's row — shared by
registration and the handover so they cannot drift apart. The holding belongs to
the assignment and goes when it does, so a seat never points at nothing running.

Tests: the resolver with and without a record, on the same and another machine,
under a former name; the store's row replaced not added, refused for an
unassigned target, removed with its assignment; CanHold's four answers and that
they follow the store. Full suite green against a real NATS and store.
2026-09-27 23:22:20 +02:00
jschoubben 1cfe6be9c4 The move ends with the old broker going, not staying
`rollout check` said the old broker stays running as an ordinary provider of
amqp, and this was not its retirement. That was ADR 0127, which ADR 0131 has
superseded: AMQP is not a provision, so once every machine reports on the new bus
nothing of the mesh speaks to the old broker and its module is unassigned. The
plan says so, as its last step.

The flag that made the "it stays" line conditional is gone with the line — there
is no case in which the broker is kept. The test that pinned the opposite now
pins this, and says which record changed under it. The stale citation of a
record numbered 0119 is corrected while here.
2026-09-27 23:04:59 +02:00
jschoubben 63ca073938 The store owns the seat set, so only the control plane may judge a claim
A claim on a seat was checked against `SeatNamed` inside `ParseManifest`, and the
build machine parses manifests too. It has no store, so there it answered from the
set compiled into the binary — a copy of data the control plane owns (ADR 0122).

When the two disagreed, that copy won where it mattered. The store's row said the
bus seat answers for `amqp`; the binary's said `mesh-bus`; and a holder that
provides `amqp` was refused at build time for not providing `mesh-bus`. The seat
went unheld, the controller lost the address it composes through that seat, and the
control plane crash-looped on a bus that was healthy the whole time.

So the two checks that read the set — a seat's scope, and what its holder must
provide — move to CatalogueProblems, which runs only in the control plane and only
after UseSeats has replaced the set with the store's. The parser keeps what it can
judge from the manifest alone, the reserved-namespace rule included.

A test pins it: the same manifest, two different values in the store, and the answer
follows the store both times. It fails if the check moves back.
2026-09-27 21:59:40 +02:00
jschoubben 497f8ea567 Merge main: the trunk renamed the seats and made them data
Both branches changed the seat set from the same starting point, so every number
collided and every `mesh-*` name existed twice. The trunk's numbers and names win:
this branch's records became 0129/0130 and its migrations 0037/0038, and the
hardcoded rename map gave way to the trunk's `seat_alias` table — a rename is a
row now (ADR 0122), not a recompile.

Three of my checks were wrong and the merge is what showed it:

A seat with an empty protocol is a marker, not an incomplete declaration. Most
node-scoped seats are markers — which module is this machine's packet filter —
and refusing one refused most of the set, the showcase module included. A
mistyped field name is already refused by the parser, so an empty protocol was
written as one deliberately.

A claim on a seat this manifest does not declare is not the parser's to judge. A
module may hold a seat another module declared; that is the whole reason ADR 0126
has callers name the seat and not its provider. Whether the seat exists is a fact
about the catalogue, so the refusal is at registration, where every declaration
is in view.

And a seat may share a name with the provision it delivers. `git`, the npm
registry and the artifact store still do, because renaming a delivering seat
cascades to every consumer requiring it, with a window where a holder stops
resolving mid-flight. The trunk deferred exactly those three on purpose.

Full suite green against a real NATS and store.
2026-09-27 18:50:18 +02:00
jschoubben dded086b54 rollout check: whether this mesh could move its bus, and what is missing
The rollout moves every node at once, so there is nothing to inspect afterwards and no
half to roll back — either the mesh was ready or it was not. That makes the readiness
question the valuable half: it costs nothing, it can be asked of a mesh that is serving as
many times as you like, and every answer is a thing somebody can go and fix.

It reads from records and dials once. Is a bus answering, does a machine hold the seat, has
that machine been sent the composed user list, does every machine have a credential for the
new bus, does every module that speaks. Each missing thing names its own next step, because
"not ready" that cannot be acted on is not an answer — and this is read at the point where
the next step is irreversible.

**A machine with no credential is the one that must stop it.** It keeps running and cannot
come back, and afterwards there is no bus to tell it anything over, so the remedy has to
happen first. The message says so.

A module that never reaches the bus is not counted as missing a credential. A third of the
catalogue never speaks, and listing those would bury the ones that matter.

`rollout --confirm` refuses and says why: the move is not being written before its check has
been run against a real mesh. And the plan it prints says the old broker stays — it remains
an ordinary provider of `amqp` for whatever else uses it, which on this installation is a
whole automation layer that has nothing to do with the mesh. This move is not its retirement,
and that is why it is survivable: what breaks if it goes wrong is the mesh's ability to
change things, not the services its modules serve.
2026-09-27 17:59:01 +02:00
jschoubben 8ceec32692 The mesh owns the operator's ~/.ssh: account fact + home-scoped resources (to-be 29)
A node carries its operator account (name + home; migration 0036, Node.Account,
SetAccount, 'node account' CLI). The account and its home are offered as
machine facts ${machine:account} / ${machine:account-home}, and machineInto
now resolves placeholders in a resource's path and owner (not just content), so
a module writes into a person's home naming what it cannot know. A RosterFile
gains Home: the file is placed under the account's home and chowned to it, its
template sees each node's Account, and a machine with no account gets none —
this is how the ssh Host blocks for every node reach a person's ~/.ssh. Roster
carries per-node accounts (Rendering.Accounts). Tested, including ssh-client
composed end-to-end. Not deployed.
2026-09-27 17:50:56 +02:00
jschoubben 5fcde512bc A build is announced under both names on the old bus, or merging breaks the live mesh
Found by asking what merging this would do to the mesh that is actually running — the
only place the question could have been asked, because the tests were green and both
buses were self-consistent.

Moving the build outcome to the role means a catalogue built from the current manifests
listens for the role's name. The catalogue *already running* listens for the module's,
because that is what it was told when it was installed. The two do not meet, so merging
as it stood would have stopped the live mesh's module graph being updated — silently,
since a binding that matches nothing is not an error.

A rename on a live bus needs the publisher and the subscriber to change together, and a
deployment cannot promise which arrives first. So the old bus announces under both names
and the order stops mattering. The module's own name retires with the bus, in step 5's
list; nothing has ever run on the bus being built, so there is no legacy name there and
this doubling has no counterpart.
2026-09-27 17:39:39 +02:00
jschoubben e5007a7daa The user list is composed before anything moves onto the bus
Found by reading the live mesh's own notes before touching it, which is where this
was heading next.

Composing the bus's user list was gated on the controller already being on the new
bus. That cannot work: the server needs its user list *before* anything moves onto
it. Step 2 of the whole change is exactly that — the server stands in the mesh
carrying nothing, on its own ports, while every node stays where it is. Under the
old gating that step was impossible: the module would come up, find no accounts
file, and its entrypoint would wait for one the controller had decided not to write.

So the only question is whether this machine runs the module that asked for the
file. A mesh that never moves has written a user list nothing reads, costing a few
hundred bytes on one node. The reverse cost a step that could not be taken.

Pinned by a test over the records of a mesh mid-change: everything running, nothing
on the new bus, and a user list that contains the controller — because a file
without it is a bus its own writer cannot connect to.
2026-09-27 17:33:52 +02:00
jschoubben cb77f35a27 A module may watch a role's events, and the catch-up turns out to be unnecessary
Moving the build outcome onto its role broke the one module that consumes it, and my own
agreement check passed anyway. The catalogue's subscription derived
`mesh.mod.mesh-build-machine.event.built` — a module namespace for a role's event, which no
such module owns — so it started, connected, and its graph stayed empty. The check compared
names, and the names agreed: the build machine does emit `built`. Only the subjects
disagreed, and a subscription that matches nothing is silence.

A consumed name is a module's event unless it names a role, and this package cannot tell by
looking — so whoever resolved the declaration says which, the way it already does for a seat
held or used. A module that watches a role gets the role's event subject and a consumer
filtered on it; watching grants subscribe and nothing else, because hearing what a role
announced is not taking part in it.

The check now compares the two halves that actually have to match — the subject a consumer
subscribes against the subject an emitter publishes — with a case pinning that it catches
this exact confusion. Comparing names was checking the easy half.

**And that answered the open question about catch-up: there is nothing to build.** The
mechanism exists because a queue on the old bus receives only what is published after it is
bound, so everything built before the catalogue existed was announced to nobody. A stream is
a log and a consumer is a position in it: a consumer created afterwards starts at the
beginning, so the builds are simply there. Asked of a real server, since the whole decision
rested on it — three builds published with nothing listening, then a consumer created, and
all three waiting for it.
2026-09-27 17:22:30 +02:00
jschoubben 47d412e13f A module declares its fail2ban jail; the mesh composes them per node (to-be 31)
The mechanism, mirroring Filtering: a module declares Jails (name, failregex,
jail stanza) naming no node/path (ADR 0112); the intrusion-prevention holder
declares Jailing (where composed jails go); the mesh gathers every assigned
module's jails into one jail.d file (a fixed id the fail2ban service restarts
on) plus a filter.d file per jail. A node not running a module has none of its
jails. Tested. Behaviour-neutral until a service module declares a jail — the
per-service content (postgres/mssql/mailu failregex+logpath) is authored next,
against how each container actually logs.
2026-09-27 17:20:50 +02:00
jschoubben 53e8f5bdd8 A person may be issued, listed and revoked
Design 25 §7's first item, which existed as a permission model and as nothing a person
could actually be given. There is a record now, and three commands.

Their authority is a list of tools and nothing else. Not a module: they hold no seat,
nothing is addressed to them, nothing is delivered to them, and they have no consumer to
acknowledge. What they have is permission to ask — which is why there is no scope and no
node in the record.

Stating what somebody may call replaces what was there rather than adding to it: a list
that could only grow is a permission nobody can take back. Forgetting somebody takes
their credential with them, because a person's row gone with their bus user left behind
is a credential that still works and that nothing derives — the worst of both, since it
keeps working and nobody can explain why.

The credential is printed once and the mesh keeps only a hash, the same contract a token
has. And it starts working at the next composition rather than immediately, because the
bus's users are a file — said out loud in both the issue and the revoke messages, since
"revoked" that still works for another minute is worth knowing about.

Four properties held by test, each a way of being wrong that would not announce itself:
a person may publish exactly the tool subjects they were given and nothing on control,
nodes or events; they cannot answer a request; changing the list removes what is no longer
named; and forgetting them revokes them.
2026-09-27 17:07:19 +02:00
jschoubben 19d2725c13 A module can name the mesh's range: ${machine:mesh-range} (novox/hq ADR 0112)
A module cannot know the private network's CIDR — it is a per-mesh value chosen
at genesis — but sometimes must name it: an intrusion filter that must never
ban a tunnel peer. Carry the overlay range on the Rendering and offer it as the
machine fact mesh-range, the same way a machine's own address is offered, so the
module names it rather than hardcoding a value (data is the mesh's). Absent when
the mesh has no range. Enables the fail2ban ignoreip fix.
2026-09-27 16:52:00 +02:00
jschoubben 8e2824201a Genesis can raise a mesh on the new bus, and the carried user list is checked against the composer
The mesh writes its own user list, and at genesis there is no mesh yet to write it. So
the installer carries the first one — the controller's own account at a well-known
bootstrap password, exactly as the store is reached at `postgres:bootstrap` and the old
bus at `guest:guest`, and rotated with them. From the controller's first composition
onward the file is the controller's.

That left a gap I would not have found by reading: the controller's own account is
created before there is a controller to mint one, so nothing recorded a hash for it, and
its first composition would have left the writer out of the file it was writing — a bus
nothing can connect to, produced by the thing connected to it. It now records a hash of
the credential it is actually using, and only if none is recorded, so a restart cannot
put the bootstrap password back over a rotated one.

The carried list and the derived one are two statements of one fact, so a test compares
them: every subject the controller derives must be in the template, and nothing wider.
It earned itself immediately — the composer was granting both a role's whole event
branch and the one event it actually follows, which is a wider way of saying the same
thing, and the wider one wins. Only the submitting half of a role is granted now; what
comes back is named exactly.

Getting this wrong is the worst kind of silent. A controller whose carried permissions
are narrower than the ones it derives comes up, connects, and is refused on the first
thing it tries, with an authorisation error naming a subject and not the template that
forgot it — and a mesh cannot be raised twice to find out.
2026-09-27 16:39:19 +02:00
jschoubben 6da9a5478b Seats keep their former names, so a rename breaks nothing (ADR 0122, phase 2)
Phase 1 made the set data; a rename still broke every reference to the old
name. This adds the stable identity: a seat's canonical name changes and its
old name becomes an alias that resolves to it forever. SeatNamed and the holder
and display matching resolve a name (former or current) to its seat, so a
manifest's claim, a held record, the git-seat lookup and the build machine's
embedded set all go on working unchanged after a rename. seat_alias table
(migration 0035), inventory Aliases/RenameSeat, openInventory loads them, and a
'seat rename <from> <to>' command does the whole thing — one operation, no
rebuild, no re-registration, no freeze. Behaviour-neutral until a seat is
renamed. Validated against postgres.
2026-09-27 16:32:22 +02:00
jschoubben 2ec0fd218b Seats are data the controller owns, loaded from its store (ADR 0122, phase 1)
The seat set was a Go slice compiled into the controller and referenced by
name everywhere, so changing it meant a rebuild and a freeze-prone deploy. It
is now a table: catalogue keeps the shipped set as defaultSeats (the seed and
the fallback) and a loadable working set; inventory adds the seat table
(migration 0034), Seats to read it, and SeedSeats to fill it idempotently
without overwriting an operator's edit; migrate seeds it; openInventory loads
it, and an empty or unreadable table leaves the compiled defaults in force so
it can never brick the control plane's boot.

Behaviour-neutral: the seeded table equals the defaults. Phase 2 (reference by
a stable id so a rename touches no manifest or code, and the builder reads the
set from the mesh) follows.
2026-09-27 16:04:44 +02:00
jschoubben e4e960ec1c A build is work submitted to a role, on both buses
ADR 0121 carried through to working code. `Builders` is the asking side and
`BuildMachine` the taking side, each with an implementation per bus, and the builder
binary and the `build` command now go through them.

On the bus being built, one publish does what two did. The old bus answered the asker
through a reply queue and announced to an events exchange, because two audiences meant
two topologies. Here the outcome is the role's own event: the asker matches it by the
id its request carried, the controller records it, the catalogue places it in the graph.
So a build machine publishes once, needs a reply queue for nothing, and needs a grant
over nobody's inbox — which is what ruled out the alternatives.

The outcome carries the module name now. Only the manifest says what was built, and on
the old bus the separate announcement carried it; with one message for three readers it
belongs in the result. A failed build names none, because it produced no module version
and the catalogue would otherwise place something that was never made.

Checked against a real server: the whole round trip; a third party on the role's event
hearing the same outcome the asker did, which is the claim the decision rests on; work
leaving the queue once settled, so no second machine repeats it; work submitted with no
machine holding the role waiting instead of failing, and being done when one arrives;
and work a machine handed back coming round again.

One thing I got wrong twice now and have written down where it bit: binding to a
consumer must name that consumer's own filter subject, not the narrower subject the
caller cares about. The client compares the two and refuses anything that is not equal,
with "subject does not match consumer".
2026-09-27 16:01:53 +02:00
jschoubben cc69737934 The agreement check knows the mesh's own roles, and was passing vacuously without them
It read the roles modules declare and not the mesh's own, so the first consumer of a
role's event was skipped as "the emitter is not installed" — which is exactly the
silence the check exists to break. It passed, and it was checking nothing.

Now it is handed the mesh's own roles too, and there is a case pinning that a consumer
of a role event the role does not emit is caught. A check that cannot fail is worse
than no check, because it reads as evidence.
2026-09-27 15:37:21 +02:00
jschoubben 0c83ecf1b5 The mesh's own roles carry a protocol, and the build branch retires
ADR 0121, first half. The `mesh-*` seats said who does a job and nothing about what
may be said to them or by them, so the mesh had roles it could not describe. They
take the same three fields a module's seat has now, and the machinery that already
derives a work queue, a holder's worker and a permission set from a declared seat
does it for these too.

The build-machine role accepts a build and emits an outcome, so `mesh.build.request`,
`mesh.control.built` and the BUILDS stream are gone. A work queue shared by several
build machines is what a seat's `accepts` already is, and keeping a second mechanism
for it was two places a permission could be wrong.

The controller's own side of a seat is a named list rather than something derived: it
is not a module and declares no `uses`, so which roles the mesh itself submits work to
has to be stated — and stating it makes that question answerable.

Two things this caught:

**The followed event subjects were hard-coded and had just gone stale.** They were
written out while the catalogue still spelled its events as the old bus's routing keys,
so converting those (issue 127) turned the pair into a controller listening to a
subject nothing publishes — the same fault as the issue, from the other side. They
derive from the emitter and the event name now, through the same function the
permission uses, so the two cannot drift apart.

**A role's queue exists before its holder**, checked against a real server, and
asserting twice changes nothing. Work queues until somebody arrives to do it, so
assigning a build machine later flushes the backlog instead of having lost it.
2026-09-27 15:34:44 +02:00