Commit Graph
90 Commits
Author SHA1 Message Date
jschoubben 0014984116 An older merge does not move a source
The forge announces what it finds merged, and an old merge surfacing late moved the recorded head
backwards and rebuilt everything built from that repository, once per old merge. A merge made
before the source was last seen is history; one that says nothing about when is taken as news.
The catalogue now carries when each source was last seen. The bus-records test follows #116:
a module's tools are every one under its own name.
2026-09-28 05:12:37 +02:00
jschoubben 35252af665 A build records the bases it was handed, and the mesh reads its edges from builds
Bases reach a recipe as build arguments, so the digest was never in the file the builder read
edges from: no build on the mesh recorded what it stood on, and 'build --on', the bases-first
order and the merge follow-up all walked a graph with no edges (novox/hq 04-ISSUES/131). The
builder now reports every base it resolved; the controller records them by artifact path and
reads the newest build's edges from the store, since a recorded manifest carries no build.on.
2026-09-28 03:51:14 +02:00
jschoubben e8aa7ed9e7 The move mints every credential and tells each machine its membership
`rollout mint` gives every principal the new bus will have a credential it does
not yet have and puts each where its owner reads it: a machine's as a membership
— bus address, fingerprint, password, transport — sealed into its declaration
(migration 0041, the `bus-membership` resource the host reads after applying); a
module's as its broker secret, through the same delivery `module issue` uses; the
control plane's own as its `bus` secret. Idempotent, and worked out from where the
bus's module is assigned rather than from this process's environment, because this
process is still on the old bus when it runs and must be.

This is the half of design 28 task 5.2 the first live attempt found missing: a
credential was minted only at enrolment, at `module issue` and for a person, so no
machine already enrolled could ever be moved. `rollout check` was right to refuse;
now there is something to run first.
2026-09-28 00:16:24 +02:00
jschoubben f325073982 AMQP is not a provision: the bus seat delivers the bus, and the word is refused
Two halves of novox/hq ADR 0131. Migration 0040 moves the mesh-broker row from
`amqp` to `mesh-bus`, so the seat's holder answers for the mesh's bus and not for
the wire protocol the old broker spoke — which is what let only the retiring
broker hold the seat that names the bus. Safe under the current holder: the
control plane composes its own address through the seat by name and the overview
derives holders by name; only registration and provision resolution read the
column. What must not happen in between is re-registering the current holder.

And the parser refuses a manifest that provides or requires `amqp`, each refusal
saying what to do instead: a module reaches the mesh's bus through the sdk and
depends on the seat, not on a protocol. A whole-catalogue test asserts nothing
beside this checkout names it; the three modules that did are removed there.

Two tests that used the old broker as a fixture now use the module that replaces
it or a manifest this package owns.
2026-09-27 23:29:00 +02:00
jschoubben 585a6abbdd A seat is handed over as one act, and the holder is on record
`seat <name> --to <node>/<module>` makes one assignment the holder of a seat in
the same write that removes the previous one. The row is new (migration 0039);
without one, the resolver derives the holder as it always did — the sole eligible
assignment, two refused — so nothing changes for a mesh that never hands a seat
over. With one, the recorded assignment holds and any other whose module could
hold the seat is eligible and silent: not refused, not holding. That is what lets
the next holder run beside the current one until the switch (hq design 26, design
28 task 5.3, ADR 0131).

Why: the controller finds its own bus through a seat, and the day that seat was
left with nobody in it — because two eligible holders could not coexist and the
old one's claim was taken away — the control plane looped for two hours while
every service stayed up. A handover that is never empty in between is the fix,
not a workaround for it.

`CanHold` is the one judgement of whether a module may hold a seat — claims it at
its scope, provides what it delivers, against the store's row — shared by
registration and the handover so they cannot drift apart. The holding belongs to
the assignment and goes when it does, so a seat never points at nothing running.

Tests: the resolver with and without a record, on the same and another machine,
under a former name; the store's row replaced not added, refused for an
unassigned target, removed with its assignment; CanHold's four answers and that
they follow the store. Full suite green against a real NATS and store.
2026-09-27 23:22:20 +02:00
jschoubben 497f8ea567 Merge main: the trunk renamed the seats and made them data
Both branches changed the seat set from the same starting point, so every number
collided and every `mesh-*` name existed twice. The trunk's numbers and names win:
this branch's records became 0129/0130 and its migrations 0037/0038, and the
hardcoded rename map gave way to the trunk's `seat_alias` table — a rename is a
row now (ADR 0122), not a recompile.

Three of my checks were wrong and the merge is what showed it:

A seat with an empty protocol is a marker, not an incomplete declaration. Most
node-scoped seats are markers — which module is this machine's packet filter —
and refusing one refused most of the set, the showcase module included. A
mistyped field name is already refused by the parser, so an empty protocol was
written as one deliberately.

A claim on a seat this manifest does not declare is not the parser's to judge. A
module may hold a seat another module declared; that is the whole reason ADR 0126
has callers name the seat and not its provider. Whether the seat exists is a fact
about the catalogue, so the refusal is at registration, where every declaration
is in view.

And a seat may share a name with the provision it delivers. `git`, the npm
registry and the artifact store still do, because renaming a delivering seat
cascades to every consumer requiring it, with a window where a holder stops
resolving mid-flight. The trunk deferred exactly those three on purpose.

Full suite green against a real NATS and store.
2026-09-27 18:50:18 +02:00
jschoubben 8ceec32692 The mesh owns the operator's ~/.ssh: account fact + home-scoped resources (to-be 29)
A node carries its operator account (name + home; migration 0036, Node.Account,
SetAccount, 'node account' CLI). The account and its home are offered as
machine facts ${machine:account} / ${machine:account-home}, and machineInto
now resolves placeholders in a resource's path and owner (not just content), so
a module writes into a person's home naming what it cannot know. A RosterFile
gains Home: the file is placed under the account's home and chowned to it, its
template sees each node's Account, and a machine with no account gets none —
this is how the ssh Host blocks for every node reach a person's ~/.ssh. Roster
carries per-node accounts (Rendering.Accounts). Tested, including ssh-client
composed end-to-end. Not deployed.
2026-09-27 17:50:56 +02:00
jschoubben cb77f35a27 A module may watch a role's events, and the catch-up turns out to be unnecessary
Moving the build outcome onto its role broke the one module that consumes it, and my own
agreement check passed anyway. The catalogue's subscription derived
`mesh.mod.mesh-build-machine.event.built` — a module namespace for a role's event, which no
such module owns — so it started, connected, and its graph stayed empty. The check compared
names, and the names agreed: the build machine does emit `built`. Only the subjects
disagreed, and a subscription that matches nothing is silence.

A consumed name is a module's event unless it names a role, and this package cannot tell by
looking — so whoever resolved the declaration says which, the way it already does for a seat
held or used. A module that watches a role gets the role's event subject and a consumer
filtered on it; watching grants subscribe and nothing else, because hearing what a role
announced is not taking part in it.

The check now compares the two halves that actually have to match — the subject a consumer
subscribes against the subject an emitter publishes — with a case pinning that it catches
this exact confusion. Comparing names was checking the easy half.

**And that answered the open question about catch-up: there is nothing to build.** The
mechanism exists because a queue on the old bus receives only what is published after it is
bound, so everything built before the catalogue existed was announced to nobody. A stream is
a log and a consumer is a position in it: a consumer created afterwards starts at the
beginning, so the builds are simply there. Asked of a real server, since the whole decision
rested on it — three builds published with nothing listening, then a consumer created, and
all three waiting for it.
2026-09-27 17:22:30 +02:00
jschoubben 53e8f5bdd8 A person may be issued, listed and revoked
Design 25 §7's first item, which existed as a permission model and as nothing a person
could actually be given. There is a record now, and three commands.

Their authority is a list of tools and nothing else. Not a module: they hold no seat,
nothing is addressed to them, nothing is delivered to them, and they have no consumer to
acknowledge. What they have is permission to ask — which is why there is no scope and no
node in the record.

Stating what somebody may call replaces what was there rather than adding to it: a list
that could only grow is a permission nobody can take back. Forgetting somebody takes
their credential with them, because a person's row gone with their bus user left behind
is a credential that still works and that nothing derives — the worst of both, since it
keeps working and nobody can explain why.

The credential is printed once and the mesh keeps only a hash, the same contract a token
has. And it starts working at the next composition rather than immediately, because the
bus's users are a file — said out loud in both the issue and the revoke messages, since
"revoked" that still works for another minute is worth knowing about.

Four properties held by test, each a way of being wrong that would not announce itself:
a person may publish exactly the tool subjects they were given and nothing on control,
nodes or events; they cannot answer a request; changing the list removes what is no longer
named; and forgetting them revokes them.
2026-09-27 17:07:19 +02:00
jschoubben 8e2824201a Genesis can raise a mesh on the new bus, and the carried user list is checked against the composer
The mesh writes its own user list, and at genesis there is no mesh yet to write it. So
the installer carries the first one — the controller's own account at a well-known
bootstrap password, exactly as the store is reached at `postgres:bootstrap` and the old
bus at `guest:guest`, and rotated with them. From the controller's first composition
onward the file is the controller's.

That left a gap I would not have found by reading: the controller's own account is
created before there is a controller to mint one, so nothing recorded a hash for it, and
its first composition would have left the writer out of the file it was writing — a bus
nothing can connect to, produced by the thing connected to it. It now records a hash of
the credential it is actually using, and only if none is recorded, so a restart cannot
put the bootstrap password back over a rotated one.

The carried list and the derived one are two statements of one fact, so a test compares
them: every subject the controller derives must be in the template, and nothing wider.
It earned itself immediately — the composer was granting both a role's whole event
branch and the one event it actually follows, which is a wider way of saying the same
thing, and the wider one wins. Only the submitting half of a role is granted now; what
comes back is named exactly.

Getting this wrong is the worst kind of silent. A controller whose carried permissions
are narrower than the ones it derives comes up, connects, and is refused on the first
thing it tries, with an authorisation error naming a subject and not the template that
forgot it — and a mesh cannot be raised twice to find out.
2026-09-27 16:39:19 +02:00
jschoubben 6da9a5478b Seats keep their former names, so a rename breaks nothing (ADR 0122, phase 2)
Phase 1 made the set data; a rename still broke every reference to the old
name. This adds the stable identity: a seat's canonical name changes and its
old name becomes an alias that resolves to it forever. SeatNamed and the holder
and display matching resolve a name (former or current) to its seat, so a
manifest's claim, a held record, the git-seat lookup and the build machine's
embedded set all go on working unchanged after a rename. seat_alias table
(migration 0035), inventory Aliases/RenameSeat, openInventory loads them, and a
'seat rename <from> <to>' command does the whole thing — one operation, no
rebuild, no re-registration, no freeze. Behaviour-neutral until a seat is
renamed. Validated against postgres.
2026-09-27 16:32:22 +02:00
jschoubben 2ec0fd218b Seats are data the controller owns, loaded from its store (ADR 0122, phase 1)
The seat set was a Go slice compiled into the controller and referenced by
name everywhere, so changing it meant a rebuild and a freeze-prone deploy. It
is now a table: catalogue keeps the shipped set as defaultSeats (the seed and
the fallback) and a loadable working set; inventory adds the seat table
(migration 0034), Seats to read it, and SeedSeats to fill it idempotently
without overwriting an operator's edit; migrate seeds it; openInventory loads
it, and an empty or unreadable table leaves the compiled defaults in force so
it can never brick the control plane's boot.

Behaviour-neutral: the seeded table equals the defaults. Phase 2 (reference by
a stable id so a rename touches no manifest or code, and the builder reads the
set from the mesh) follows.
2026-09-27 16:04:44 +02:00
jschoubben 0c83ecf1b5 The mesh's own roles carry a protocol, and the build branch retires
ADR 0121, first half. The `mesh-*` seats said who does a job and nothing about what
may be said to them or by them, so the mesh had roles it could not describe. They
take the same three fields a module's seat has now, and the machinery that already
derives a work queue, a holder's worker and a permission set from a declared seat
does it for these too.

The build-machine role accepts a build and emits an outcome, so `mesh.build.request`,
`mesh.control.built` and the BUILDS stream are gone. A work queue shared by several
build machines is what a seat's `accepts` already is, and keeping a second mechanism
for it was two places a permission could be wrong.

The controller's own side of a seat is a named list rather than something derived: it
is not a module and declares no `uses`, so which roles the mesh itself submits work to
has to be stated — and stating it makes that question answerable.

Two things this caught:

**The followed event subjects were hard-coded and had just gone stale.** They were
written out while the catalogue still spelled its events as the old bus's routing keys,
so converting those (issue 127) turned the pair into a controller listening to a
subject nothing publishes — the same fault as the issue, from the other side. They
derive from the emitter and the event name now, through the same function the
permission uses, so the two cannot drift apart.

**A role's queue exists before its holder**, checked against a real server, and
asserting twice changes nothing. Work queues until somebody arrives to do it, so
assigning a build machine later flushes the backlog instead of having lost it.
2026-09-27 15:34:44 +02:00
jschoubben 05ff6065d0 Event names are checked now, per manifest and across the catalogue
Issue 127 stood because nothing compared the two halves. Every manifest was
well-formed on its own and every derivation correct on its own, and no
cross-module subscription in the mesh matched anything — a subscription that
matches nothing is not an error, it is silence.

Two checks, because the mistake is possible at two scales.

Per manifest: an event is a local name, and `module.` is refused with the name to
write instead. A module emitting under what reads as another module's name is
refused too, pointing at the seat, where a name outlives whoever holds it.

Across the catalogue: where a consumed event's emitter is present, it must emit
that event. It cannot demand a live emitter for everything — a module lives in its
own repository and may be installed long before the one whose events it wants — so
the rule is narrower and still catches this. It found two real dangling
subscriptions the moment it ran.

Wildcards were undecided and two manifests needed them: `*` is one name and `**`
is the rest, spelled the mesh's way and derived to `>` here and `#` on the old bus.
A manifest naming either would stop being true when the wire changed, which is the
whole reason names are local.

And the field documentation taught the old form, examples included — which is why
the drift was uniform across 37 manifests rather than scattered. Nobody was
guessing; everybody followed the comment.
2026-09-27 14:43:16 +02:00
jschoubben ee1b8ffe24 1.7, second half: the user list read out of the mesh's records
The derivation had nothing feeding it. `BusRecords` reads what it needs — the
machines, what each runs, every manifest, and which machines hold a live token —
and turns it into the records the composer derives from.

**A module's authority comes from its manifest, not from its assignment.** The
assignment says where it runs; what it may say is what it declared. So the two are
read together and the manifest decides, which is also why a seat's protocol is
gathered across the whole catalogue rather than from one manifest: a seat is
declared by one module and held by another, and that is the whole reason a seat
exists.

Three things checked against a real store, each a user that would be wrong in a
way nothing reports:

- A module assigned to a machine becomes a user with exactly the authority it
  declared, including the protocol of a seat some *other* module declared — a
  module granted nothing on a seat it was assigned to send to would fail on its
  first publish with an authorisation error that says nothing about a seat.
- Only a machine holding a live token gets an enrolment user. One outliving its
  token is a right to join that nobody issued.
- A module assigned and absent from the catalogue is refused rather than composed
  with an empty permission list. The catalogue already refuses to forget an
  assigned module, so this is the second line — and it earns its place there,
  because relying on another package's invariant is how a rule ends up enforced by
  nothing.

People are left empty rather than guessed at: the account model is built and
`operator issue` is not, so there is nobody to derive yet.
2026-09-27 01:49:09 +02:00
jschoubben 0560c792d8 1.7, first half: the mesh can say who its bus users are, and hold their keys
Two pieces the composer has been waiting for since it was written.

**The credential has to outlive its own minting.** On the bus the mesh runs on
today an account is a management call: mint a password, hand it over, seal the
plaintext to whoever will use it, keep nothing — which works because the broker
remembers. Here the users are one file, rewritten whenever any of it changes, so
keeping nothing would mean the first person's access change silently blanking
every module's password. So a bus user's bcrypt hash is now recorded, keyed by the
username the file needs, and the plaintext comes back exactly once. Verified
against a real store that the hash verifies the password it was made from, that
the password itself is not in there, that minting again rotates rather than adds,
and that forgetting a node takes its host's and its modules' credentials with it.

**Permissions are not stored, and that is the point.** Only the credential is
kept. Authority is derived from what each module declares, every time the file is
written (ADR 0043) — a stored permission list would be a second account of a
user's authority, able to disagree with the records it came from, and both would
look internally consistent while they did.

`Users` derives the list: the controller always first and always present, one
user per node, one per module per node, one per live token, one per person. Two
users with one name is refused where both can be named, rather than left to be
whichever one the server happened to read. A user the mesh has never minted a
password for is *named* rather than dropped or written as a user anybody is:
that is an ordinary situation with an obvious remedy, and the caller decides
whether a partial file is worth writing.

What remains of 1.7: delivering the file to the node that runs the server, and
minting at enrolment and assignment — which is transport-coupled, because a node
on the old bus must not be handed a credential for the new one.
2026-09-27 01:44:53 +02:00
jschoubben 06cf3c04e5 The consume side behind a seam, and the window wiring into the loop
The outbound half went behind `Bus` and the transport stopped reaching its
callers; this is the other half, and the larger one. Every handler took
`amqp.Delivery`, so the serving loop could not move to another bus without
moving enrolment, reports, builds, upgrades and catch-up with it in one breath.

`Control` states one message in the mesh's words — took it, dropped it, or held
it for the store — and `Inbound` is where messages come from. The AMQP
implementation is today's loop moved rather than changed: same queues, same
prefetch, same holding, because the mesh is running on it and a bus nothing
speaks yet is no reason to alter the one every node is on.

The window (window.go) is now what decides, instead of the conditions that were
inlined in the loop. Two things that surfaced in the wiring:

**Supersession is asked before the store, not after.** A report about a
declaration the mesh has moved past would otherwise wait out a restarting store
to be written and then overwrite what the node is doing now.

**Half of a report is not about a declaration, and that half is never stale.**
What the machine *is* — the tunnel it took over, the ports its own bundle
holds, what an adopted node found, a node moving its overlay key — reaches the
mesh on a report and nowhere else. A rekey set aside as stale is a node whose
overlay key never moves, and no retry is coming, because the node said it once.
So staleness is asked only of a report that is purely an apply's account.

The one thing holding-in-memory can do that holding-in-the-server cannot is
named rather than hidden: `About` sets aside a held message when a newer one
about the same thing arrives, and the bus being built ignores it because the
digest answers the same question.
2026-09-27 00:44:16 +02:00
jschoubben 7fc5fd02fd The mesh's interface takes over the found tunnel's MTU
Carries MTU from the reported tunnel (mesh-host#28) through inventory,
the overlay graph's TakeOver, into the generated config's [Interface].
A tuned path keeps its MTU across the takeover instead of regressing to
1420 and hanging transfers no ping would reveal. Two emit tests; a
tunnel with no MTU writes no line.
2026-09-26 22:40:42 +02:00
jschoubben 952092ccb3 A carried peer is nameable, and the mesh answers for it (hq 112)
The tunnel the hub took over routes to machines the predecessor knows
by name and the mesh knew only by address — taking the resolver in that
state silences three machines at once. Now the operator states which
machine a carried address is (overlay name <address> <name>), the
statement rides tunnel_peer.named, and namesInTheMesh answers for named
not-yet-enrolled peers — one reading, so the hosts fact, a container's
hosts and the resolver cannot disagree. Enrolment verifies the word:
a machine enrolling under a named peer's key with a different name is
refused where the operator can read it, the stated name keeps the
carried address, and an enrolled peer's name is the node's — naming it
again refuses. The issue's rule holds: a name the predecessor answers
for keeps resolving until the machine behind it is a node.
2026-09-26 20:09:42 +02:00
jschoubben 50734095b8 A repeat assignment says nothing changed (ADR 0115)
One assignment of a module per node is now the rule, not a limitation —
the operator dropped the multi-assignment requirement, and the schema's
(node, module) key has been the decision since migration 0005. What
changed: Assign reports whether the assignment was new, and the command
says 'already runs — one node runs one of each (ADR 0115); nothing
changed' instead of printing 'is assigned' for a no-op, which read as
an action that happened. Idempotence stays: a repeat is exit 0, because
a script stating what is already true is not wrong.
2026-09-26 19:02:35 +02:00
jochen 97448194ac Seats are a closed set, a seat's holder answers for what it delivers, and a build source may live on the git seat
Implements novox/hq ADR 0110 and 0111.

The seat set lives in internal/catalogue/seats.go: fourteen seats, each with a scope, what occupying
it delivers, and the record that made it one. A test asserts the count and a decision per entry, so
changing the set means finding the argument, as the host's vocabulary test does. The first set is
every seat already claimed — including the-private-network, which the network module claims from a
manifest composed in this repository's code, not from any module.json — plus npm-package-registry
(ADR 0109) and git (ADR 0111). A test parses every catalogue manifest and this repository's own and
fails on any refused claim, so closing the set refuses nothing in use.

ParseManifest now refuses a claim on a seat the mesh does not define, a seat claimed at another
scope, and a delivering seat claimed by a module that does not provide what it delivers. A
malformed claim is refused once, for being malformed.

Resolution: among several providers of a mesh provision, a pin still wins; then the holder of the
seat that delivers it; then the only provider; otherwise refused as before. ADR 0009's "never
guessed" holds — the seat is the choice made once, mesh-wide, rather than a pin per consumer node.
A provider now carries the module it came from, because a provider is a (node, module) pair and the
pair is what tells a holder from a neighbour on the same machine.

The planner's second pass is now given the first pass's holdings. Without them, a node consuming a
seat-delivered provision was refused there, and a refused node's own claims dropped out of what the
mesh holds — letting a second holder of one of its seats pass unrefused.

`seats [--json]` lists every seat, what it delivers, and each holder, derived from assignments
every time and never stored. Unheld seats are listed. A stored claim outside the set — possible
for a manifest registered before the set closed, since stored manifests are not re-validated — is
shown rather than hidden.

`build --self <owner>/<repo>` builds from a repository on the git seat's holder. The clone URL is
composed at build time from the holder's node and what it serves for git; the recorded source is the
path and the seat (migration 0032), never an address, so a moved forge changes nothing recorded.
Nobody holding the seat refuses self-hosted builds and says so; external URLs are unchanged. An
address passed with --self is refused rather than recorded as a path.

Replaces three foundation tests that defended the builder's carried package binding. The catalogue
removed that binding when the builder began requiring the registry through a real grant, so the
tests were already failing on main; they now assert the builder requires what the npm seat delivers
and carries no copy of its own, and that the forge holds the npm and git seats.

Verified: go vet clean; the whole suite passes against a throwaway Postgres (make postgres), the new
inventory tests included; gofmt clean apart from cmd/mesh-builder/stdout_test.go, which fails on
main too.
2026-09-25 20:48:10 +02:00
jschoubben 4566c5c9aa Adopt the tunnel as a mesh fact, refuse a mismatched takeover, and rekey after enrolment
Review of the ADR 0105 build (hq ADR 0105). Four things it got wrong and one
path it lacked:

- A predecessor spoke's tunnel names one peer, the hub, routed the whole
  range; recording refused it and the whole enrolment failed. Range-routed
  peers are skipped now — only the hub's peers are ever carried.
- The range and the carried peers were conditions on the node being adopted,
  so converging the hub would have renumbered the mesh and dropped the peers
  still reaching it. They are facts of the tunnel record now, mode aside; the
  takeover alone is declared to an adopted node. Converging the hub is refused
  while a carried peer has not enrolled, naming it.
- A push composed a takeover for a hub whose address or endpoint disagreed
  with the tunnel, which would have the host stop the found interface and
  raise the mesh's where no peer listens. The graph refuses to compose it,
  naming both and the placement that fixes it.
- The host's account said taken or not; "found down and the mesh's not up"
  read as not taken. Three states now, and an account on every takeover.
- A hub that enrolled before this feature holds a key of its own, and
  re-enrolling would rotate every key the mesh sealed credentials to. A node
  now rekeys in a report, signed with its identity key over the key it
  leaves, the key it takes and the tunnel; the mesh verifies against the live
  key, refuses a stale or foreign proof, records key and tunnel, and moves a
  hub to the tunnel's address. `overlay show` names the path for a hub that
  found no tunnel.

Also: a carried IPv6 peer is routed /128, and identity.ForTest exists so the
link can be tested against a real identity store.
2026-09-24 00:02:07 +02:00
jschoubben 3c836f0abb Adopt the predecessor's tunnel in place: its range, its address, its peers
On an adopted hub the private network takes over the tunnel it finds rather
than running beside it (hq ADR 0105): two tunnels leave the mesh's unreachable
through the provider's filter, so no machine can ever join.

The node presents the found tunnel when it enrols, under the key it took as
its own; the inventory records it (node.tunnel, tunnel_peer — migration 0031)
and the mesh composes from it: the overlay's range is the adopted tunnel's,
the hub is placed at the tunnel's address on the tunnel's port, and every
peer the tunnel had is carried in the hub's peer list as a peer of the
tunnel, not a node of the mesh, until a node enrols with that key — which
then keeps the address the tunnel had for it. A fresh node never gets an
address the tunnel holds. The hub's declaration tells the host which unit to
take over; the host's account of carrying it is recorded and shown.

Every reader of the range follows the setting; nothing stores it. A found
tunnel under another key is recorded and not adopted, so ADR 0100's
non-overlap rule keeps applying where a tunnel is left running beside the
mesh's. A lab bed and test skeleton for "How it is checked" are under lab/.
2026-09-23 23:26:34 +02:00
jschoubben 2cf8739a84 Clear the whole account an adopted node gave when it converges (hq ADR 0100) 2026-09-22 19:45:35 +02:00
jschoubben 87ecc9326e Refuse a port given for the whole mesh where it is set, not at every node's composition (hq ADR 0100) 2026-09-22 19:45:35 +02:00
jschoubben 65187d4de0 Wait for a held node without pinning a pool connection, and give up after a bounded wait naming it (hq ADR 0100) 2026-09-22 18:31:10 +02:00
jschoubben 689c2c6d33 Hold a node from composing to sending, so a push composed before converge or adopt is never sent after it (hq ADR 0100) 2026-09-22 18:10:04 +02:00
jschoubben 4bb19c9e40 Give a machine port one holder: refuse ssh's, another module's and a doubled one, and release the assignment a given port replaces (hq ADR 0100) 2026-09-22 18:05:31 +02:00
jschoubben dbf62b5212 Read the foundation's ports from each node's settings, wherever a port is used (hq ADR 0100) 2026-09-22 17:31:07 +02:00
jschoubben 28894fa5bd Keep what an adopted node reports holding, the firewall it found and what is reachable on it (hq ADR 0100) 2026-09-22 17:23:09 +02:00
jschoubben 1c32af6a22 Record whether a node is adopted and which modules were taken on it (hq ADR 0100) 2026-09-22 17:14:05 +02:00
jschoubben 4567fa666c Review of 083: finishing an enrolment whose token was spent takes proof of the key's private half, a live lease and a first delivery — a public key alone cannot replay a spent token; shutdown leaves held messages for the broker; identical builds supersede; what is held leaves room in the prefetch 2026-09-22 14:33:13 +02:00
jschoubben a3b7e830c8 Review of 083: a message the store cannot take is held and retried on a ticker, not slept on, so enrolments are answered meanwhile; a newer one per subject supersedes; the password is replaced after the spend; the same presenter may finish after a lost answer; upgrades retry only on the store 2026-09-22 14:23:03 +02:00
jschoubben 1a41b88ed3 Nothing the control queue carries is lost while the store restarts: an enrolment claims its token and spends it last, and is asked to try again; build results, upgrades and catch-ups are handed back, bounded (novox/hq issue 083) 2026-09-22 14:06:28 +02:00
jschoubben 3fd0c37d22 Review of 082: a store killed mid-conversation is an outage and a wrong password is not; one report holds the queue at most two minutes, then is let go loudly; shutdown does not wait out the pause 2026-09-22 13:34:04 +02:00
jschoubben 047830a6b5 A report the store cannot take yet is handed back to the broker, not acknowledged and lost (novox/hq issue 082) 2026-09-22 13:25:13 +02:00
jschoubben 0d1f2a4152 Review: module issue's pre-check factored and tested; a pair delivery for a requirement the module keeps no secret for is refused; builder issue's usage says the module comes first 2026-09-21 23:58:36 +02:00
jschoubben 35c5c2bb9b secret accept is refused for a name the module does not declare, and a pair delivery for a requirement or local it has not got (novox/hq issue 078) 2026-09-21 23:31:06 +02:00
jschoubben 9f3790dcda Review: one local name is still a local name; a local name is unique; recovery knows it; recipes read as instructions; a tag before a digest; ask fails at once when nothing serves
A secrets object with one local name delivered no file. Two requirements could share
a local name. secret recover and the export could not tell two locals apart. The
recipe check missed continued lines and read heredoc bodies as bases. repo:tag@digest
kept the tag in the repository. ask now publishes mandatory, so a tool nothing serves
is said at once rather than after the wait.
2026-09-21 21:03:22 +02:00
jschoubben 6ae4ae1dba A module may hold several secrets from one provider, each a pair of its own
secrets: maps a requirement to several files under local names. Each local name is
its own need, its own pair credential (the pair is keyed on it: migration 0027),
its own file on the consumer, its own holder at the provider (the identity with the
local name after it) and rotates apart from the others. The plain shape is
unchanged and every existing row is the credential it was (novox/hq 04-ISSUES/069,
ADR 0094).
2026-09-21 20:28:16 +02:00
jschoubben 53a79a17a1 Review: a failure is the same by resource id, not by the host's words; bound files and /run/docker.sock are declared
The host's error text may carry a duration or a counter, and a resource looping on
it would never have read as stuck. The previous row is read and compared here.
Stuck needs a start to say. A container may mount the file a binding lands in; the
runtime socket is declared under both of its spellings; the catalogue-wide test
takes MESH_CATALOG.
2026-09-21 19:23:02 +02:00
jschoubben 396e05bb65 An operator delivers a pair credential, and the mesh never replaces it
secret accept grows --provider: the value is sealed to the consumer's node, the
provider's node and the operator's key, and the pair records origin 'accepted'.
An accepted pair is not remade when a key changes (the mesh does not hold the
value; the read is refused naming the remedy) and rotate refuses it (accepting a
new value is the rotation). The vault's third species has its entry
(novox/hq 04-ISSUES/070, ADR 0092).
2026-09-21 17:50:41 +02:00
jschoubben fcaad7271a A failure that repeats is said to be stuck
The mesh kept one report per machine, replaced, so a resource nothing can ever apply
looked like a failure that had just happened, every reconcile interval, for ever.
The row now keeps when the current failure began and how many reports in a row have
said it — the same outcome, refusal and failed resources; anything different starts
again and a clean apply clears it. Three make the machine stuck, and status says so
beside the failure, in words and in JSON (novox/hq 04-ISSUES/065, ADR 0090).
2026-09-21 17:43:24 +02:00
jschoubben 77e6c1a684 Recoverable means sealed to the current operator key; recovery names the provider
From review: the export counted any operator-sealed row as recoverable, so a
secret sealed to a replaced key was reported as openable with the current one;
replacing the key counted orphans in one table of two; and a pair credential
held from two providers was recovered as whichever row came first. The export
now lists what the current key opens, what an earlier key opens, and what has
no copy; `secret recover` takes --provider and refuses ambiguity; files that
must not exist are created exclusively; one constructor builds the export for
the operator's file and the vault's disk alike.
2026-09-21 01:16:32 +02:00
jschoubben 565f144a20 A pair credential is sealed to the operator key too
The secret the vault provides a module is the credential of the consumer↔vault
pair, and so is every credential a provider grants; sealing only own secrets
to the operator left exactly those unrecoverable. Same column, same call; the
export and `secret recover` address a pair by consumer node, module and the
provision's name, and say which kind each entry is.
2026-09-21 00:36:16 +02:00
jschoubben e140ed5d0b An operator key, a second seal on every own secret, and the vault keeps the export
novox/hq ADR 0085, amended: the mesh's root secrets — the store's superuser,
the broker's administrator, every secret a module holds for itself — were
sealed to a node key and nothing else, so a lost node took them with it.
Now the mesh records an operator's public sealing key and seals every own
secret to it as well, minted or accepted. The private half is written once
by `operator key new` to a file the operator keeps off the mesh; the mesh
holds one more blob per secret that it cannot open.

`secret recover` opens a secret with that key, to a 0600 file, from the
store or from an export; `secret export` writes every operator-sealed copy
as ciphertext. A module that `keeps` (the vault) is handed that export as a
declared file on its own disk, so recovery survives the store.

Secrets made before the key exists have no operator copy and are said so —
the plaintext was discarded — until each is issued again.
2026-09-20 23:56:49 +02:00
jschoubben c3b88b9148 Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben 3ae7b88c6d A catalogue asks for what it was not there to hear
Its event queue is durable, so a running catalogue misses nothing. What it cannot
have is what was announced before it first ran — and on a fresh mesh that is never
arbitrary: the shared base, the store the catalogue runs on, and the catalogue
itself are each necessarily built BEFORE a catalogue exists to hear about them.
The graph's foundation is the part it never sees.

So it says it is catching up, and the control plane re-announces what it
recorded, oldest first, marked as a replay. Oldest first because a graph is built
in the order things happened: registering a module that stands on a base before
the base would point an edge at a version nothing has seen, and the shape of a
fresh mesh guarantees the base is both first and the one that was missed.

The replayer hands announcements back rather than publishing them, because the
wire belongs to the link package and a replay building its own events could drift
from what the builder emits — the one thing it must match exactly, since the
catalogue has a single handler for both.

Its own queue and its own consumer: two consumers on one queue split its
messages, and a catch-up request going to whichever half was not listening is a
gap that looks like a working mesh.

Toward novox/hq 04-ISSUES/050.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 01:26:58 +02:00
jschoubben 2b82872ac3 A build keeps what it was told
The builder announces a build with the resolved manifest, the path inside the
repository, and every artifact it stood on. The control plane received all of it
and kept none of it.

That was survivable while the catalogue heard the same announcement directly. It
stops being survivable the moment the catalogue was not there to hear it — which
on a fresh mesh is always, and always for the same modules: the shared base, the
store the catalogue runs on, and the catalogue itself are each necessarily built
BEFORE the catalogue exists to hear about them. The graph's foundation is the
part the graph never sees.

Replaying those builds needs what they said, not a summary. Without the manifest
there are no requires/provides edges; without `against` there are no build edges,
which are the ones that answer "a base moved, what must be rebuilt". A replay
carrying neither would restore the module list and leave the question the
catalogue exists for still wrong, while looking fixed.

Kept null rather than empty where a build predates this, so a replay can say it
is holding nothing instead of inventing an empty declaration for a module that
certainly had one. And `built_against`, not `built_on`: that column exists and
means the machine, which is a different fact about a different subject.

Toward novox/hq 04-ISSUES/050.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 01:22:34 +02:00
jschoubben 25d2fe1308 Only a thing built and never run is unassignable
The check read "declares no resources" as "runs nowhere", and those are not
the same. The private network declares no resources either — the control plane
computes them when it composes a machine's declaration — and it is assigned to
every machine that has to reach another one. Refusing it stopped a four-machine
bed at its first assignment.

The signal is narrower: it builds an artifact and places nothing. Made a
function of its own, because a judgement with a wrong answer this expensive
should be testable without a database — nothing guarded it, which is how it
shipped.
2026-09-14 17:32:43 +02:00