Commit Graph
190 Commits
Author SHA1 Message Date
jschoubben aecac5bda2 One bus: the AMQP transport is gone from the controller
The mesh runs on the seat's bus alone (novox/hq ADR 0131, design 28 task 5.5). The old
transport's consume loop, build request, tool ask, management API and account scoping are
deleted, and the bus switch with them; the controller connects to the broker seat and to
nothing else. The store-window tests keep their assertions on a bus-less fake, and the tests
that only made sense for the old transport's in-memory holding go with it.
2026-09-28 03:36:16 +02:00
jschoubben 525f10b858 A merge on the forge builds what it moved, bases first
The controller follows the forge's merges (novox/hq 04-ISSUES/131). For each
module recorded as built from that repository and branch it records the move to
the merge commit and builds it — bases first, because a module built before the
module it stands on is built against the old one and reports success, and a base
that fails stops what stands on it. Nothing is pushed here: what a finished build
does to the machines running the module stays the upgrade's decision.

Two more things the same ordering gives: `build --behind` builds bases first, and
`build --on <module>` rebuilds everything that stands on a module — the rebuild a
changed base needs, which "behind" does not see because their sources did not
move.
2026-09-28 02:59:02 +02:00
jschoubben 9fe9b5349c Merge pull request 'A store row keeps its seat's protocol, and a holder may take work from its queue' (#108) from fix/store-seats-keep-their-protocol into main 2026-09-28 02:21:24 +02:00
jschoubben 964285f08c The build machine takes work on the bus its credential names, and the work queue has a taker
Two halves of one gap the first build over the new bus met. The machine decided
its bus from a variable its container never received, so the credential the mesh
sealed to it went unread; a credential for the new bus names the bus by scheme and
carries user, password and fingerprint beside the address, and that is enough to
dial it, pinned. And the roles' work queues were raised with no holders, so the
consumer a machine binds to take work was never created: the holders are read
from the catalogue and the handover record, as the resolver reads them.
2026-09-28 01:59:20 +02:00
jschoubben 6005a8471f A principal may hear what its consumer delivers
A push consumer delivers on _DELIVER.<its name>, and a client bound to it
subscribes exactly that. No principal was granted it, and the server refused
every one the first time it bound a consumer: the control plane, each machine,
and a module would have been next. Each kind is granted its own consumers'
delivery subjects and no other's. The line announcing the raised bus printed the
URL with the credential in it; the address alone now.
2026-09-28 01:46:16 +02:00
jschoubben 1fd6194ff8 The control plane pins the bus's certificate, and keeps its password out of errors
The bus presents the mesh's own certificate, which names nothing a public verifier
accepts; the client verified by name and failed against a bus that was answering
("certificate is not valid for any names", 2026-09-28). It now pins the leaf's
fingerprint from MESH_BROKER_CERTIFICATE, as every host does. And a connection
error named the whole URL, password included — the address alone now.
2026-09-28 01:35:20 +02:00
jschoubben 3907ea0db0 The control plane serves and pushes on the bus it is told to
The seams were there and nothing chose a side: serve, push, ask and build all
opened the old bus's connection and declared over its channel, whatever
MESH_BUS_NATS said. So the switch moved every host and left the control plane
unable to follow — "this control plane has no MESH_BROKER_AMQP" with the new bus
named and standing (2026-09-28). That was task 4.3 of design 28, still open.

One place now decides: connectLink reads the switch, refuses both buses named at
once, raises the new bus's streams and this controller's consumers when it is
handed the inventory, and opens the link over whichever bus it is on. Every
caller that sent a declaration or asked a tool through the old channel goes
through the server's bus instead, which the new transport has and the channel is
not. OverNats is that outbound: a declaration is a JetStream publish into the
node's own subject, an event is announced on the subject its name derives to, a
tool is request and reply on the module's tool subject.
2026-09-28 01:29:44 +02:00
jschoubben 4d62e6caf1 The mint leaves the control plane's old-bus secret alone
The control plane is a module too, and its broker secret is the old bus's
credential it is still using while the mint runs. Writing the new bus's blob there
cut the mesh off from its own old bus mid-move. Its new-bus credential is the
controller principal's bus secret; the module principal is skipped.
2026-09-28 01:08:11 +02:00
jschoubben 9b715524a2 The network map resolves each machine with the seat holders on record
Without them, a machine running the next holder of a seat beside the current one
resolves as two holders, is refused, and drops out of the map — and with it the
address every other machine composes for what it offers. Found live: the control
node vanished from the private network the moment the new bus was assigned beside
the old one, and nothing on any machine could be composed.
2026-09-28 00:55:39 +02:00
jschoubben 7b02feaebb The mint names the bus by bare host, and can mint again
BareAddress adds a scheme where none was, so the host it yielded carried one and
every URL built from it carried two. Caught before a push: the host is now taken
with no scheme and no port, and every URL adds its own. `rollout mint --again`
mints every credential afresh for exactly this case — a mint that was wrong before
anything received it.
2026-09-28 00:37:10 +02:00
jschoubben 4b209d944d The mint finds the bus on the hub
The map of who is where on the private network lists the machines placed around
the hub, not the hub — and the control node is the hub, and runs the new bus.
Found on the first live run: refused for having no address. The address every
machine already dials the current bus at is the same machine, so that host with
the new port is what they are told.
2026-09-28 00:34:17 +02:00
jschoubben e8aa7ed9e7 The move mints every credential and tells each machine its membership
`rollout mint` gives every principal the new bus will have a credential it does
not yet have and puts each where its owner reads it: a machine's as a membership
— bus address, fingerprint, password, transport — sealed into its declaration
(migration 0041, the `bus-membership` resource the host reads after applying); a
module's as its broker secret, through the same delivery `module issue` uses; the
control plane's own as its `bus` secret. Idempotent, and worked out from where the
bus's module is assigned rather than from this process's environment, because this
process is still on the old bus when it runs and must be.

This is the half of design 28 task 5.2 the first live attempt found missing: a
credential was minted only at enrolment, at `module issue` and for a person, so no
machine already enrolled could ever be moved. `rollout check` was right to refuse;
now there is something to run first.
2026-09-28 00:16:24 +02:00
jschoubben 4c41628b20 The first handover records the standing holder without re-judging it
On a mesh that predates the record, every handover has to begin by writing down who
already holds the seat — otherwise the next holder cannot be assigned beside it,
because two eligible claimants with nothing on record are refused. Found on the
live mesh minutes after 0040 moved the bus seat's row: the standing holder no
longer satisfies what the seat delivers, on purpose, and so could not be recorded,
and so nothing could stand beside it.

Recording who already holds is not making a new holder. Derivation never read
what the seat delivers, so the standing holder holds regardless; when nothing is
on record and the named assignment claims the seat at its scope, only that claim
is checked. Every change of holder is still judged in full.
2026-09-27 23:34:40 +02:00
jschoubben 33c4e4be34 Merge pull request 'A seat is handed over as one act, and the holder is on record' (#91) from feat/seat-handover into main 2026-09-27 21:22:50 +00:00
jschoubben 585a6abbdd A seat is handed over as one act, and the holder is on record
`seat <name> --to <node>/<module>` makes one assignment the holder of a seat in
the same write that removes the previous one. The row is new (migration 0039);
without one, the resolver derives the holder as it always did — the sole eligible
assignment, two refused — so nothing changes for a mesh that never hands a seat
over. With one, the recorded assignment holds and any other whose module could
hold the seat is eligible and silent: not refused, not holding. That is what lets
the next holder run beside the current one until the switch (hq design 26, design
28 task 5.3, ADR 0131).

Why: the controller finds its own bus through a seat, and the day that seat was
left with nobody in it — because two eligible holders could not coexist and the
old one's claim was taken away — the control plane looped for two hours while
every service stayed up. A handover that is never empty in between is the fix,
not a workaround for it.

`CanHold` is the one judgement of whether a module may hold a seat — claims it at
its scope, provides what it delivers, against the store's row — shared by
registration and the handover so they cannot drift apart. The holding belongs to
the assignment and goes when it does, so a seat never points at nothing running.

Tests: the resolver with and without a record, on the same and another machine,
under a former name; the store's row replaced not added, refused for an
unassigned target, removed with its assignment; CanHold's four answers and that
they follow the store. Full suite green against a real NATS and store.
2026-09-27 23:22:20 +02:00
jschoubben 1cfe6be9c4 The move ends with the old broker going, not staying
`rollout check` said the old broker stays running as an ordinary provider of
amqp, and this was not its retirement. That was ADR 0127, which ADR 0131 has
superseded: AMQP is not a provision, so once every machine reports on the new bus
nothing of the mesh speaks to the old broker and its module is unassigned. The
plan says so, as its last step.

The flag that made the "it stays" line conditional is gone with the line — there
is no case in which the broker is kept. The test that pinned the opposite now
pins this, and says which record changed under it. The stale citation of a
record numbered 0119 is corrected while here.
2026-09-27 23:04:59 +02:00
jschoubben 497f8ea567 Merge main: the trunk renamed the seats and made them data
Both branches changed the seat set from the same starting point, so every number
collided and every `mesh-*` name existed twice. The trunk's numbers and names win:
this branch's records became 0129/0130 and its migrations 0037/0038, and the
hardcoded rename map gave way to the trunk's `seat_alias` table — a rename is a
row now (ADR 0122), not a recompile.

Three of my checks were wrong and the merge is what showed it:

A seat with an empty protocol is a marker, not an incomplete declaration. Most
node-scoped seats are markers — which module is this machine's packet filter —
and refusing one refused most of the set, the showcase module included. A
mistyped field name is already refused by the parser, so an empty protocol was
written as one deliberately.

A claim on a seat this manifest does not declare is not the parser's to judge. A
module may hold a seat another module declared; that is the whole reason ADR 0126
has callers name the seat and not its provider. Whether the seat exists is a fact
about the catalogue, so the refusal is at registration, where every declaration
is in view.

And a seat may share a name with the provision it delivers. `git`, the npm
registry and the artifact store still do, because renaming a delivering seat
cascades to every consumer requiring it, with a window where a holder stops
resolving mid-flight. The trunk deferred exactly those three on purpose.

Full suite green against a real NATS and store.
2026-09-27 18:50:18 +02:00
jschoubben dded086b54 rollout check: whether this mesh could move its bus, and what is missing
The rollout moves every node at once, so there is nothing to inspect afterwards and no
half to roll back — either the mesh was ready or it was not. That makes the readiness
question the valuable half: it costs nothing, it can be asked of a mesh that is serving as
many times as you like, and every answer is a thing somebody can go and fix.

It reads from records and dials once. Is a bus answering, does a machine hold the seat, has
that machine been sent the composed user list, does every machine have a credential for the
new bus, does every module that speaks. Each missing thing names its own next step, because
"not ready" that cannot be acted on is not an answer — and this is read at the point where
the next step is irreversible.

**A machine with no credential is the one that must stop it.** It keeps running and cannot
come back, and afterwards there is no bus to tell it anything over, so the remedy has to
happen first. The message says so.

A module that never reaches the bus is not counted as missing a credential. A third of the
catalogue never speaks, and listing those would bury the ones that matter.

`rollout --confirm` refuses and says why: the move is not being written before its check has
been run against a real mesh. And the plan it prints says the old broker stays — it remains
an ordinary provider of `amqp` for whatever else uses it, which on this installation is a
whole automation layer that has nothing to do with the mesh. This move is not its retirement,
and that is why it is survivable: what breaks if it goes wrong is the mesh's ability to
change things, not the services its modules serve.
2026-09-27 17:59:01 +02:00
jschoubben 8ceec32692 The mesh owns the operator's ~/.ssh: account fact + home-scoped resources (to-be 29)
A node carries its operator account (name + home; migration 0036, Node.Account,
SetAccount, 'node account' CLI). The account and its home are offered as
machine facts ${machine:account} / ${machine:account-home}, and machineInto
now resolves placeholders in a resource's path and owner (not just content), so
a module writes into a person's home naming what it cannot know. A RosterFile
gains Home: the file is placed under the account's home and chowned to it, its
template sees each node's Account, and a machine with no account gets none —
this is how the ssh Host blocks for every node reach a person's ~/.ssh. Roster
carries per-node accounts (Rendering.Accounts). Tested, including ssh-client
composed end-to-end. Not deployed.
2026-09-27 17:50:56 +02:00
jschoubben e5007a7daa The user list is composed before anything moves onto the bus
Found by reading the live mesh's own notes before touching it, which is where this
was heading next.

Composing the bus's user list was gated on the controller already being on the new
bus. That cannot work: the server needs its user list *before* anything moves onto
it. Step 2 of the whole change is exactly that — the server stands in the mesh
carrying nothing, on its own ports, while every node stays where it is. Under the
old gating that step was impossible: the module would come up, find no accounts
file, and its entrypoint would wait for one the controller had decided not to write.

So the only question is whether this machine runs the module that asked for the
file. A mesh that never moves has written a user list nothing reads, costing a few
hundred bytes on one node. The reverse cost a step that could not be taken.

Pinned by a test over the records of a mesh mid-change: everything running, nothing
on the new bus, and a user list that contains the controller — because a file
without it is a bus its own writer cannot connect to.
2026-09-27 17:33:52 +02:00
jschoubben 53e8f5bdd8 A person may be issued, listed and revoked
Design 25 §7's first item, which existed as a permission model and as nothing a person
could actually be given. There is a record now, and three commands.

Their authority is a list of tools and nothing else. Not a module: they hold no seat,
nothing is addressed to them, nothing is delivered to them, and they have no consumer to
acknowledge. What they have is permission to ask — which is why there is no scope and no
node in the record.

Stating what somebody may call replaces what was there rather than adding to it: a list
that could only grow is a permission nobody can take back. Forgetting somebody takes
their credential with them, because a person's row gone with their bus user left behind
is a credential that still works and that nothing derives — the worst of both, since it
keeps working and nobody can explain why.

The credential is printed once and the mesh keeps only a hash, the same contract a token
has. And it starts working at the next composition rather than immediately, because the
bus's users are a file — said out loud in both the issue and the revoke messages, since
"revoked" that still works for another minute is worth knowing about.

Four properties held by test, each a way of being wrong that would not announce itself:
a person may publish exactly the tool subjects they were given and nothing on control,
nodes or events; they cannot answer a request; changing the list removes what is no longer
named; and forgetting them revokes them.
2026-09-27 17:07:19 +02:00
jschoubben 19d2725c13 A module can name the mesh's range: ${machine:mesh-range} (novox/hq ADR 0112)
A module cannot know the private network's CIDR — it is a per-mesh value chosen
at genesis — but sometimes must name it: an intrusion filter that must never
ban a tunnel peer. Carry the overlay range on the Rendering and offer it as the
machine fact mesh-range, the same way a machine's own address is offered, so the
module names it rather than hardcoding a value (data is the mesh's). Absent when
the mesh has no range. Enables the fail2ban ignoreip fix.
2026-09-27 16:52:00 +02:00
jschoubben 8e2824201a Genesis can raise a mesh on the new bus, and the carried user list is checked against the composer
The mesh writes its own user list, and at genesis there is no mesh yet to write it. So
the installer carries the first one — the controller's own account at a well-known
bootstrap password, exactly as the store is reached at `postgres:bootstrap` and the old
bus at `guest:guest`, and rotated with them. From the controller's first composition
onward the file is the controller's.

That left a gap I would not have found by reading: the controller's own account is
created before there is a controller to mint one, so nothing recorded a hash for it, and
its first composition would have left the writer out of the file it was writing — a bus
nothing can connect to, produced by the thing connected to it. It now records a hash of
the credential it is actually using, and only if none is recorded, so a restart cannot
put the bootstrap password back over a rotated one.

The carried list and the derived one are two statements of one fact, so a test compares
them: every subject the controller derives must be in the template, and nothing wider.
It earned itself immediately — the composer was granting both a role's whole event
branch and the one event it actually follows, which is a wider way of saying the same
thing, and the wider one wins. Only the submitting half of a role is granted now; what
comes back is named exactly.

Getting this wrong is the worst kind of silent. A controller whose carried permissions
are narrower than the ones it derives comes up, connects, and is refused on the first
thing it tries, with an authorisation error naming a subject and not the template that
forgot it — and a mesh cannot be raised twice to find out.
2026-09-27 16:39:19 +02:00
jschoubben 6da9a5478b Seats keep their former names, so a rename breaks nothing (ADR 0122, phase 2)
Phase 1 made the set data; a rename still broke every reference to the old
name. This adds the stable identity: a seat's canonical name changes and its
old name becomes an alias that resolves to it forever. SeatNamed and the holder
and display matching resolve a name (former or current) to its seat, so a
manifest's claim, a held record, the git-seat lookup and the build machine's
embedded set all go on working unchanged after a rename. seat_alias table
(migration 0035), inventory Aliases/RenameSeat, openInventory loads them, and a
'seat rename <from> <to>' command does the whole thing — one operation, no
rebuild, no re-registration, no freeze. Behaviour-neutral until a seat is
renamed. Validated against postgres.
2026-09-27 16:32:22 +02:00
jschoubben 2ec0fd218b Seats are data the controller owns, loaded from its store (ADR 0122, phase 1)
The seat set was a Go slice compiled into the controller and referenced by
name everywhere, so changing it meant a rebuild and a freeze-prone deploy. It
is now a table: catalogue keeps the shipped set as defaultSeats (the seed and
the fallback) and a loadable working set; inventory adds the seat table
(migration 0034), Seats to read it, and SeedSeats to fill it idempotently
without overwriting an operator's edit; migrate seeds it; openInventory loads
it, and an empty or unreadable table leaves the compiled defaults in force so
it can never brick the control plane's boot.

Behaviour-neutral: the seeded table equals the defaults. Phase 2 (reference by
a stable id so a rename touches no manifest or code, and the builder reads the
set from the mesh) follows.
2026-09-27 16:04:44 +02:00
jschoubben e4e960ec1c A build is work submitted to a role, on both buses
ADR 0121 carried through to working code. `Builders` is the asking side and
`BuildMachine` the taking side, each with an implementation per bus, and the builder
binary and the `build` command now go through them.

On the bus being built, one publish does what two did. The old bus answered the asker
through a reply queue and announced to an events exchange, because two audiences meant
two topologies. Here the outcome is the role's own event: the asker matches it by the
id its request carried, the controller records it, the catalogue places it in the graph.
So a build machine publishes once, needs a reply queue for nothing, and needs a grant
over nobody's inbox — which is what ruled out the alternatives.

The outcome carries the module name now. Only the manifest says what was built, and on
the old bus the separate announcement carried it; with one message for three readers it
belongs in the result. A failed build names none, because it produced no module version
and the catalogue would otherwise place something that was never made.

Checked against a real server: the whole round trip; a third party on the role's event
hearing the same outcome the asker did, which is the claim the decision rests on; work
leaving the queue once settled, so no second machine repeats it; work submitted with no
machine holding the role waiting instead of failing, and being done when one arrives;
and work a machine handed back coming round again.

One thing I got wrong twice now and have written down where it bit: binding to a
consumer must name that consumer's own filter subject, not the narrower subject the
caller cares about. The client compares the two and refuses anything that is not equal,
with "subject does not match consumer".
2026-09-27 16:01:53 +02:00
jschoubben 0c83ecf1b5 The mesh's own roles carry a protocol, and the build branch retires
ADR 0121, first half. The `mesh-*` seats said who does a job and nothing about what
may be said to them or by them, so the mesh had roles it could not describe. They
take the same three fields a module's seat has now, and the machinery that already
derives a work queue, a holder's worker and a permission set from a declared seat
does it for these too.

The build-machine role accepts a build and emits an outcome, so `mesh.build.request`,
`mesh.control.built` and the BUILDS stream are gone. A work queue shared by several
build machines is what a seat's `accepts` already is, and keeping a second mechanism
for it was two places a permission could be wrong.

The controller's own side of a seat is a named list rather than something derived: it
is not a module and declares no `uses`, so which roles the mesh itself submits work to
has to be stated — and stating it makes that question answerable.

Two things this caught:

**The followed event subjects were hard-coded and had just gone stale.** They were
written out while the catalogue still spelled its events as the old bus's routing keys,
so converting those (issue 127) turned the pair into a controller listening to a
subject nothing publishes — the same fault as the issue, from the other side. They
derive from the emitter and the event name now, through the same function the
permission uses, so the two cannot drift apart.

**A role's queue exists before its holder**, checked against a real server, and
asserting twice changes nothing. Work queues until somebody arrives to do it, so
assigning a build machine later flushes the backlog instead of having lost it.
2026-09-27 15:34:44 +02:00
jschoubben 1c56210530 Name system seats by scope; let a module define its own (ADR 0121)
System seats are mesh-* (one, mesh-wide) or node-* (one per node). Renamed:
the-build-machine -> mesh-build-machine (+scope mesh), the-catalogue ->
mesh-catalog, the-dns-port -> node-dns-resolver, the-intrusion-prevention ->
node-intrusion-prevention, the-packet-filter -> node-packet-filter,
the-resolver-configuration -> node-resolver-config, the-uplink -> node-uplink.
Removed the-showcase from the set — it becomes the first module-defined seat.

A manifest may declare its own seats (DefinesSeats); a claim is a system seat,
a reserved mesh-*/node-* name the mesh does not define (refused), or a
module-defined seat valid only when the manifest declares it.

Deferred: the delivering registry seats (git, npm-package-registry,
the-artifact-store) and the-private-network (a scope + server/client model
change), per ADR 0121.
2026-09-27 14:30:56 +02:00
jschoubben eb72ec36ba 1.7 finished: minting, the file delivered, and a test flake I caused
**First, a correction: the previous commit went in on a false check.** Its message
says the suite passed; it did not. The check piped `go test` through a filter that
swallowed the failures and then printed "green" regardless. Two tests were failing
when 4de10e3 landed.

What was failing was my own doing. Purging the streams instead of deleting them
(4de10e3) left the *consumers* behind, because deleting a stream takes its
consumers with it and purging does not. A durable push consumer surviving between
tests keeps pushing to a delivery subject the previous test's subscription has gone
from: the messages count as delivered, go nowhere, and the next test waits out its
timeout for an announcement the server believes it already sent. Consumers are now
removed with the purge. Five consecutive clean runs.

`-p 1` stays, because two packages asserting and deleting the same fixed-name
objects on one bus is a real race — but its comment said the cause I had guessed
and not the one I found, so it now says the right thing.

**And delivery was not finished when I said it was.** Nothing filled
`Rendering.BusUsers`, so the composed file would never have reached a node.
`composeBusUsers` closes it: composed per push for the machine holding
`mesh-broker`, never kept, because the list is a function of the mesh's records and
a stored copy could disagree with them while both looked consistent. A user with no
credential is left out and named rather than written as a user without a password —
an ordinary situation with an obvious remedy — but a file with no users at all is
refused, because that bus would refuse every connection in the mesh.

**Minting, on both halves.** A node at enrolment and a module at `module issue`.
Three things differ from a management call and each is the point of the move: the
credential is minted into the mesh's records and becomes usable at the next
composition, so no server need be reachable; the password travels beside the address
rather than inside it, because a credential embedded in a URL leaks into every log
line that prints a connection; and a module's durable consumer is derived from what
it declared rather than named, so it cannot ask for delivery of something it did not
say it consumes.

A node reconnecting may be refused until that composition reaches the machine
running the bus. That is what the host's reconnect backoff is for and it is
survivable by design; waiting for the push would hold an enrolment open for as long
as a declaration takes to apply.

Tested that the switch is a switch: a node enrolling on one bus comes away with a
credential for that bus and none for the other, because one that held both could be
half-moved and nothing would say which half.
2026-09-27 03:19:41 +02:00
jschoubben 4de10e32e3 The bus's objects are raised on every start, and one switch says which bus
Two of 1.7's three remaining pieces.

**Raised on every start, not created once at genesis.** A stream somebody deleted,
a mesh raised from a restored backup, or a bus whose data directory was replaced
all have records and no objects — and a node whose consumer is missing hears
nothing while everything else about it looks correct.

The order is not a preference: a consumer on a stream that does not exist is
refused *naming the stream*, so somebody reading that refusal goes looking for a
deletion instead of a reversed pair of lines. Pinned by a test, along with the one
thing about seats that reads like an omission and is not — a seat's work queue is
asserted whether or not anybody holds it, because work queues until a holder
appears, so installing the module a week later flushes the backlog instead of
having lost it.

Against a real server: every object accepted, asserting twice changes nothing (a
start that failed the second time is a controller that cannot restart), a machine
joining an already-raised bus is accepted, each node's consumer is bound to its own
declaration subject and no other's, and CONTROL does not dead-letter — because the
store window's bound belongs to the controller and a server that gave up first
would discard the push the stream exists to protect.

**Which bus this mesh is on is one fact, read in one place.** Every seam the change
went behind ships both implementations; this is what the rollout flips. Being told
about both is refused at start rather than warned about: a mesh half on each is one
where a declaration goes out on one bus and the report comes back on the other, and
every component logs success while it happens — ADR 0074's failure arriving through
configuration instead of through code. The refusal names both variables and says
which to unset, because whoever reads it has to choose and the wrong choice is a
rollout half done.
2026-09-27 02:59:05 +02:00
jochen 63ba2d178f review: hold the hosts region through composition, and say so in plan
A declaration-level test composes the shipped networking module with a
resolver and asserts /etc/hosts arrives as mesh-wireguard.fact-node-names
with into: block and region-only content, that the resolver's restart-on
still names it, and that a resource's at passes through untouched — a
composition step dropping into would otherwise go unnoticed. plan --show
marks files written into, so a region is not read as the whole file.
The rollout order is spelled out: every node's host, the controller's
own included, must be block-aware before this controller ships (hq 128).
2026-09-27 00:05:41 +02:00
jschoubben 6c12780abe Describe the bus on its own terms
Comments framed the new bus by what it replaces — a comparison in almost
every explanation, which reads as though NATS were a variant of the old
thing rather than the mesh's nervous system. Removed throughout, and
OverAMQP becomes OverCurrent: the seam's two sides are the bus the mesh
runs on today and the one being built, not two protocols.

What remains is the client library's own package name, which is its name.
2026-09-26 23:51:00 +02:00
jschoubben 2fad32767e The controller's outbound link behind a seam, with both transports
Step 3.4, first half. Every one of these took an *amqp.Channel, so the
transport reached every caller and swapping it meant touching all of them.
The seam turned out to be small — the controller sends exactly two kinds of
message that expect no answer — which is the same measurement that said
this bus could be replaced at all.

Bus is stated in the mesh's words, not a transport's: PublishEvent and
PublishDeclaration. Two implementations, both shipping, because steps 1 to
4 leave every node on AMQP and the NATS one is selected at the rollout.
Both ship is also what makes them comparable: one conformance fixture holds
both to the same envelope, and the NATS one is checked against a real
server reading back from the stream rather than from the code that wrote it.

Still on *amqp.Channel: RequestBuild and Ask, which carry reply-queue
machinery, and the whole consume side — the control loop, enrolment, serve.
2026-09-26 23:47:15 +02:00
jschoubben e14b02991e The controller marks a deliberately-empty declaration owns_nothing (hq 127)
The host refuses an empty body unless told the emptiness is meant
(mesh-host#29). When a node's declaration composes to no resources —
which #77 now sends rather than skips — Body() sets owns_nothing, so
the node applies it and drops what it last held. A declaration with
resources never carries the marker. One test.
2026-09-26 23:40:04 +02:00
jschoubben 0a8a592ef4 An empty declaration is sent, so a node drops what it last held (hq 127)
push skipped any node whose declaration composed to zero resources. A
node that HELD something before — the broker opening a placement gave
an adopted node, say — then kept it forever: the empty declaration that
would drop it was never sent, and the node's own heartbeat re-applied
the stale resource with no way for the mesh to say it is gone. Now the
empty declaration is sent; the host drops what the mesh owned and keeps
what it found. A node that never held anything applies it as a no-op.
Surfaced on ace: the foundation-opening fix (#74) removed its only
resource, and the correction could not reach it until this.
2026-09-26 23:32:30 +02:00
jschoubben 173c8c7c21 Rename the mesh's seats to mesh-*, keeping their interfaces
novox/hq ADR 0118: the prefix is the reservation rule, so a module
declaring any mesh-* name is refused and there is no reserved-names list to
drift. Ten seats renamed in the table, the manifests that claim them, the
controller's own shipped manifests, and the tests.

Not the migration 0118 expected: a holding is derived at resolution from
manifests and never stored, so nothing recorded points at an old name. A
kept rename table tells a manifest written against one what it became —
kept rather than retired, because a module lives in its own repository and
may be registered long after the catalogue stopped using it.

**A seat is not the interface it delivers.** The git seat became mesh-git
and the git provision did not; likewise the package registry. A blanket
replace renamed both, and the failure read "the package registry is served
on <nil>", which does not say "you renamed an interface". A test now pins
every seat against what it delivers, and that neither name is also the
other.
2026-09-26 23:07:09 +02:00
jschoubben 7fc5fd02fd The mesh's interface takes over the found tunnel's MTU
Carries MTU from the reported tunnel (mesh-host#28) through inventory,
the overlay graph's TakeOver, into the generated config's [Interface].
A tuned path keeps its MTU across the takeover instead of regressing to
1420 and hanging transfers no ping would reveal. Two emit tests; a
tunnel with no MTU writes no line.
2026-09-26 22:40:42 +02:00
jschoubben cc252472e2 A taken tunnel brings its ListenPort, even on a node the hub cannot dial
A home node behind NAT (no Endpoint → not Reachable) that took over a
tunnel must still listen on that tunnel's port: its LAN peers dial it
there. ListenPort was gated on Reachable, which conflated 'a peer dials
me here' with 'the hub can dial me' — so the takeover guard refused
overlay-up, and the guard's suggested remedy (re-place with an
endpoint) breaks a NAT'd node's path: it stops keepalive and hands the
hub a private LAN address to dial. TakeOver now carries the found
tunnel's port (already known to the controller), and the interface
listens on it when the node is not otherwise reachable. Two tests;
Endpoint-reachable nodes keep the old path unchanged.
2026-09-26 22:33:27 +02:00
jschoubben 48d8c89749 The broker opening is only on the broker's host, not every node
Enrolling ace applied adoption.opening-tcp-5671-incoming to it, opening
5671 from anywhere (v4+v6) where nothing listens — the ace session
caught it. foundation ports widen the broker's from:mesh port to
from-anywhere so a machine that is not yet on the mesh can make its
first dial; that belongs on the broker's host alone. foundationPortsFor
keeps the port only when a module resolved onto this node listens on
it, so novox opens 5671 and a node that merely dials out opens nothing.
Two tests, both directions.
2026-09-26 22:20:10 +02:00
jschoubben 952092ccb3 A carried peer is nameable, and the mesh answers for it (hq 112)
The tunnel the hub took over routes to machines the predecessor knows
by name and the mesh knew only by address — taking the resolver in that
state silences three machines at once. Now the operator states which
machine a carried address is (overlay name <address> <name>), the
statement rides tunnel_peer.named, and namesInTheMesh answers for named
not-yet-enrolled peers — one reading, so the hosts fact, a container's
hosts and the resolver cannot disagree. Enrolment verifies the word:
a machine enrolling under a named peer's key with a different name is
refused where the operator can read it, the stated name keeps the
carried address, and an enrolled peer's name is the node's — naming it
again refuses. The issue's rule holds: a name the predecessor answers
for keeps resolving until the machine behind it is a node.
2026-09-26 20:09:42 +02:00
jschoubben 50734095b8 A repeat assignment says nothing changed (ADR 0115)
One assignment of a module per node is now the rule, not a limitation —
the operator dropped the multi-assignment requirement, and the schema's
(node, module) key has been the decision since migration 0005. What
changed: Assign reports whether the assignment was new, and the command
says 'already runs — one node runs one of each (ADR 0115); nothing
changed' instead of printing 'is assigned' for a no-op, which read as
an action that happened. Idempotence stays: a repeat is exit 0, because
a script stating what is already true is not wrong.
2026-09-26 19:02:35 +02:00
jschoubben 7e42380dcd Merge main 2026-09-26 14:29:00 +02:00
jschoubben 6ac9013d6e builder: a clone may offer the forge's credential, through git's own store
A private repository could not be built: the builder clones anonymously,
and had no way to say who it is. It already holds exactly one credential
to exactly the right place — the package-registry binding and its sealed
secret, one gitea user whose password answers npm and git alike — so a
clone now offers that, and nothing new is minted or carried.

Offered, never pushed: the credential is written as a git
credential-store file (0600, in the workspace, never argv) and named
with -c credential.helper, so git itself decides when it applies — only
on an authentication challenge, and only for the URL it was written
for, scheme, host and port included. A public repository clones exactly
as before; a repository on any other host is never shown it. The same
store rides along on an artifact's own context clone, so a private
module with a private context builds too.
2026-09-25 21:47:32 +02:00
jochen 97448194ac Seats are a closed set, a seat's holder answers for what it delivers, and a build source may live on the git seat
Implements novox/hq ADR 0110 and 0111.

The seat set lives in internal/catalogue/seats.go: fourteen seats, each with a scope, what occupying
it delivers, and the record that made it one. A test asserts the count and a decision per entry, so
changing the set means finding the argument, as the host's vocabulary test does. The first set is
every seat already claimed — including the-private-network, which the network module claims from a
manifest composed in this repository's code, not from any module.json — plus npm-package-registry
(ADR 0109) and git (ADR 0111). A test parses every catalogue manifest and this repository's own and
fails on any refused claim, so closing the set refuses nothing in use.

ParseManifest now refuses a claim on a seat the mesh does not define, a seat claimed at another
scope, and a delivering seat claimed by a module that does not provide what it delivers. A
malformed claim is refused once, for being malformed.

Resolution: among several providers of a mesh provision, a pin still wins; then the holder of the
seat that delivers it; then the only provider; otherwise refused as before. ADR 0009's "never
guessed" holds — the seat is the choice made once, mesh-wide, rather than a pin per consumer node.
A provider now carries the module it came from, because a provider is a (node, module) pair and the
pair is what tells a holder from a neighbour on the same machine.

The planner's second pass is now given the first pass's holdings. Without them, a node consuming a
seat-delivered provision was refused there, and a refused node's own claims dropped out of what the
mesh holds — letting a second holder of one of its seats pass unrefused.

`seats [--json]` lists every seat, what it delivers, and each holder, derived from assignments
every time and never stored. Unheld seats are listed. A stored claim outside the set — possible
for a manifest registered before the set closed, since stored manifests are not re-validated — is
shown rather than hidden.

`build --self <owner>/<repo>` builds from a repository on the git seat's holder. The clone URL is
composed at build time from the holder's node and what it serves for git; the recorded source is the
path and the seat (migration 0032), never an address, so a moved forge changes nothing recorded.
Nobody holding the seat refuses self-hosted builds and says so; external URLs are unchanged. An
address passed with --self is refused rather than recorded as a path.

Replaces three foundation tests that defended the builder's carried package binding. The catalogue
removed that binding when the builder began requiring the registry through a real grant, so the
tests were already failing on main; they now assert the builder requires what the npm seat delivers
and carries no copy of its own, and that the forge holds the npm and git seats.

Verified: go vet clean; the whole suite passes against a throwaway Postgres (make postgres), the new
inventory tests included; gofmt clean apart from cmd/mesh-builder/stdout_test.go, which fails on
main too.
2026-09-25 20:48:10 +02:00
jschoubben 524cc2a3ec plan: show what a module would open and why
Every module.json already declares a why for each port under listens,
but plan only ever used it to build the firewall's rule set — nothing
printed it. An operator deciding whether to assign a module had no way
to see what it would open without reading the manifest by hand.

plan <node> now prints each assigned module's listens entries — port,
protocol, source, and its why — right under the module line, so the
same text that feeds the firewall is visible at the point someone is
actually deciding whether to open it.
2026-09-24 18:40:30 +02:00
jschoubben 2277583e99 Tell the resolver the machines, not the names the mesh merely serves
The map the control plane hands a resolution holds both: the machines, and every name
the mesh was told to route to whichever machine serves it. A container's hosts wants all
of it, so a routed name resolves to the proxy. A resolver's zones want only the machines:
told the mesh's suffix is its own it answers authoritatively for everything under it and
forwards none of it, so a routed name with the suffix appended — drive.example.test.internal
— is a name nobody will ever ask for, standing beside the machines and looking as real.

Found composing the resolver's first assignment on a live machine, before pushing it.
hq issue 111.
2026-09-24 01:31:11 +02:00
jschoubben 0d8264ff55 Give the resolver the mesh's suffix as a local domain and a module its machine's address
hal dnsmasq-app conversion, hq 08-connectivity. Converting the resolver from the module it
replaces made it forward what it cannot answer, which is what the predecessor's does, and
that found two things the controller did not say.

A resolver that forwards must not send a mesh name it does not know upstream: the
`node-zones` fact now carries `local=/<suffix>/` beside the wildcards, written here rather
than in the daemon's configuration because the suffix is the mesh's choice and this file is
the one place the mesh writes what it chose. The default lives in one helper now instead of
being spelled in two functions.

The predecessor points the container runtime's `dns` at the machine's own tunnel address —
a container cannot reach the machine's loopback. A module writing that key needs the
address, and `${machine:at}` is the machine's name; a runtime's resolver list cannot be a
name it would need that resolver to look up. So a module may say `${machine:address}`: what
`at` resolves to, read from the same names the hosts file and the wildcards are written
from, absent — and refused — off the network like `at` is.

The `mesh-resolver` and `resolver-data` constants go: nothing provided or consumed either,
the fact and `mesh-addressing` are the mechanism, and a requirement nothing provides is
refused at resolution.

Tests: the catalogue's dnsmasq, resolv-conf and resolved-split-dns manifests are parsed
and composed as a machine would receive them — fixed upstreams, no-resolv, 127.0.0.1, the
machines file, the runtime's key, the pair that decides what a machine asks refused on one
node; and on a real mesh the resolver's machines file is composed with a wildcard per
machine on the network and composed again without one that left, mirroring the hosts fact.
2026-09-24 01:10:15 +02:00
jschoubben 6073e94a4f Merge pull request 'Adopt the predecessor's tunnel in place: its range, its address, its peers (hq ADR 0105)' (#49) from feat/adopt-the-tunnel into main 2026-09-23 22:38:31 +00:00
jschoubben 4566c5c9aa Adopt the tunnel as a mesh fact, refuse a mismatched takeover, and rekey after enrolment
Review of the ADR 0105 build (hq ADR 0105). Four things it got wrong and one
path it lacked:

- A predecessor spoke's tunnel names one peer, the hub, routed the whole
  range; recording refused it and the whole enrolment failed. Range-routed
  peers are skipped now — only the hub's peers are ever carried.
- The range and the carried peers were conditions on the node being adopted,
  so converging the hub would have renumbered the mesh and dropped the peers
  still reaching it. They are facts of the tunnel record now, mode aside; the
  takeover alone is declared to an adopted node. Converging the hub is refused
  while a carried peer has not enrolled, naming it.
- A push composed a takeover for a hub whose address or endpoint disagreed
  with the tunnel, which would have the host stop the found interface and
  raise the mesh's where no peer listens. The graph refuses to compose it,
  naming both and the placement that fixes it.
- The host's account said taken or not; "found down and the mesh's not up"
  read as not taken. Three states now, and an account on every takeover.
- A hub that enrolled before this feature holds a key of its own, and
  re-enrolling would rotate every key the mesh sealed credentials to. A node
  now rekeys in a report, signed with its identity key over the key it
  leaves, the key it takes and the tunnel; the mesh verifies against the live
  key, refuses a stale or foreign proof, records key and tunnel, and moves a
  hub to the tunnel's address. `overlay show` names the path for a hub that
  found no tunnel.

Also: a carried IPv6 peer is routed /128, and identity.ForTest exists so the
link can be tested against a real identity store.
2026-09-24 00:02:07 +02:00
jschoubben e07b56ce43 An address is read from the node's settings where it is used, never recorded with a port
Three readers did not follow a moved foundation port (novox/hq 04-ISSUES/102),
and each took the control-node down in its own way: the control plane's own
store and broker connections, sealed at genesis with the port inside; and every
build the mesh ever recorded, kept as `<registry>:<port>/<module>/<artifact>@…`.

The control plane cannot open its own sealed connections to move a port, and it
cannot bind the store as a consumer would — a binding mints a credential. So its
settings get a third twin, `NAME_PORT`, read on top of the sealed value by the
store, the broker, the management API and the bus connection, and filled into
its container by a placeholder that names a seat, `${seat:mesh-store:5432}`,
from the node's given or mesh-assigned ports — never the manifest's number, and
empty when the mesh has nothing to add, so what genesis wrote stands. A value
that is still a placeholder is nothing said, aloud: the manifest naming it lands
in the next commit, once every control plane that composes it knows it.

A build is now recorded by digest and path — `artifact-store://<module>/<artifact>@…`
— and the store's address is composed in where a reference is used: the
declaration, the trust file, the bases a build is handed, a replay to the
catalogue. Over the network as `<node>.internal:<port>`; on the store's own node
before any network exists — every genesis push before its "network" step — by
loopback. A reference recorded before this, with an address, is re-routed the
same way when the mesh built it. The trust file and every provider's address
come from one derivation: the node's given port, over the mesh's assignment,
over the manifest's number.

novox/hq 04-ISSUES/102
2026-09-23 23:49:31 +02:00