Commit Graph
102 Commits
Author SHA1 Message Date
jochen 8ddc019cd2 Grant the self-check its ban-list question, say a refusal at once, judge the engine by its delivered version (hq to-be 45 Phase 1)
Live on 2026-10-06, two of the first self-check's findings were its own:

- D8 asked every machine's node-intrusion-prevention.banned, and the
  controller's grant did not name the subject: the bus refused it 24 times
  and D8 timed out after thirty seconds instead of saying so. The verbs the
  self-check asks are named in broker.VerbsTheSelfCheckAsks and granted
  (mesh.seat.<seat>.tool.<verb>.*); each probe declares the seat verbs it
  calls, askSeatTool refuses an undeclared one, and a test over the
  registry fails a probe whose question the controller is not granted.
  AskSeatTool now returns a refused publish at once ("the bus refused…")
  instead of waiting out its timeout; D8 asks the machines in parallel.
- D10 read every machine as behind right after a push: a node-engine says
  its version as the directory it is delivered into, the archive's digest
  (31045596c83a, catalogue versionOf), and D10 compared that with the
  build's commit (1545b00a). It now compares with the versions the
  registered build is delivered as, and a hand-placed engine's commit.
2026-10-06 10:44:38 +02:00
jochen bb1607e424 Say when the mesh is wrong: conditions, watchdogs, the bus's advisories, doctor (hq to-be 45 Phase 1)
Every one of the 48 core failures of research 031 was found by a person
looking; the mesh's answers carried the fact for whoever asked and told
nobody.

- The condition store (to-be 45 §2): mesh-controller_conditions, one key
  per open condition, written by compare-and-set so a person's silence
  and the watchdogs never lose each other's word; every transition kept
  ninety days in mesh-controller_condition-history and said as the
  seat's events condition-raised / condition-changed / condition-cleared
  (the condition at the top level, with event, at, change, why, show),
  offered again while the bus is away. Raised and cleared by observation
  only; a clearing reopened within ten minutes is the same condition with
  its count up, its silence kept. Verbs: conditions, conditions show,
  conditions silence (a hand act, at most a week), conditions history.
- ADR 0224's provider standing is the first kind, provider-failing, held
  by the provider's events; the provider_standing table is no longer read
  or written (left in place: dropping it is the operator's word).
- status leads with the open conditions, urgent first, and says all well
  only with none open; conditions it cannot read are said and not well.
- The signals table compiled in, one watchdog loop over it every 30s: S1
  heartbeat (3 intervals, asleep machines excepted, control node urgent
  after 30 min), S2 report after a send, S3 plan tier, S4 event loop deaf,
  S5 merge not acted, S6 ask lost, S7 call hung, S8 provider silent, S9
  advisories, S10 self-check silent, S11 node tools silent, S13 stale
  refusals; S12, S14, S15 deferred with their reasons. A row that cannot
  see raises probe-failed and clears nothing. A test generated from the
  table suppresses each signal inside and past its bound.
- The bus's advisories (maximum deliveries, a mesh consumer deleted) and
  the controller's own slow consumer and refused subjects, said in the
  mesh's words.
- doctor: the probe registry D1-D10 (D5 deferred) and DW, every five
  minutes, each in thirty seconds; a probe that cannot run is never a
  pass. D1 validates with mesh-host's own validator. Every run ends with
  the doctor-heartbeat event mesh-watcher listens for.
- The controller is granted its new buckets, events, the two advisories
  and $SRV.INFO; the node tools their tools-alive heartbeat. The streams
  and consumers the controller asserts and the ones D6/D7 expect are one
  derivation.
2026-10-06 10:21:11 +02:00
jochen e74c32ed50 Keep calls and hand acts on the bus, answer status at once, record durations (hq to-be 45 Phase 0)
A controller restart lost every call's outcome, `status` composed the mesh
while its caller waited (18.6s live on 2026-10-06, past the 10s window), a
repair by hand left no trace, and the core's bounds had nothing measured to
be set from.

- calls: kept in the controller's bucket mesh-controller_calls (last 1000 or
  14 days, answers bounded to 64 KiB), read by id across a restart; a
  controller starting marks a stopped one's running calls abandoned; each
  call names its caller from the inbox its answer goes to.
- status: the serving controller composes it at start, after news from a
  machine, a build or an acting verb, and every minute; the verb answers the
  last composition at once with when and how long it took. Composing resolves
  each machine once instead of twice.
- hand-act log in mesh-controller_hand-acts: push (required through the seat),
  plans stop/close, broker consumer-reset and the new hand-act record take
  --why/--cause/--condition; `hand-acts` lists them and repeated causes;
  status counts the week's.
- durations (migration 0066): apply (send to first report), heartbeat gap,
  plan tier and build, recorded as heard; `durations` summarises them.
- the controller's seat row takes this binary's definition of its own verbs,
  so the console no longer judges calls against an older build's schema.
- the controller is granted its two buckets' subjects.
2026-10-06 02:59:36 +02:00
jochen 09c0c6b367 Keep the account of the sent declaration over an older one (hq issue 267)
The last report stored per node decides whether a release plan moves on,
and it was whichever arrived last. A report about a declaration the mesh
has moved past now records what it says about the machine but leaves the
account of the apply alone, so arrival order cannot undo the newer.
2026-10-06 01:45:35 +02:00
mesh-admin d1fc25f682 Merge pull request 'Catch up on merges the bus announced and never handed over (hq issue 266)' (#73) from fix/missed-merges-are-caught-up into main 2026-10-05 23:33:18 +00:00
jochen 59f4d486b1 Catch up on merges the bus announced and never handed over (hq issue 266)
The controller acted only on what its events consumer handed it, so a merge
the bus skipped left modules behind with nothing said. The stream is now read
back every five minutes on a single-filter consumer, and any merge that would
still move a module after ten minutes is said and acted on.
2026-10-06 01:29:21 +02:00
jochen 801552c0eb Answer every seat call within ten seconds and keep what came of it (hq issue 265)
A push outlasted the console's 30s wait and, when it sent the bus its
changed user list, the broker's reload forgot the reply it may send:
the push happened and its caller was told it did not answer. Calls now
answer in full or as running with an id, a push answers before it
sends, refused answers are recorded on their call, and 'calls' reads
them back.
2026-10-06 01:14:58 +02:00
jochen 8d9d33ae85 Report a provider that keeps failing a consumer in status (hq ADR 0224)
The identity provider failed every consumer for a day and status called the
mesh well (hq issue 179). The controller now follows every provider's
provisioner.failing/recovered, keeps the newest failing word per provider,
machine and consumer (migration 0065), and status, its JSON and node show
name it until it recovers. Every module that receives contributions is
granted the two events, so no manifest can forget them.
2026-10-06 00:13:23 +02:00
jochen 2421b82ad2 Keep a named push from sending builds a policy or a plan holds back
A named push flushed every other machine whose declaration differed from
what it was last sent (hq ADR 0083). Under an upgrade policy of `record`,
or a plan still waiting on its first machine (ADR 0218), every machine
running the module differs, so `push <one>` sent the held build to all of
them (hq issue 259).

Each send now records which build of each module it carried
(node.sent_builds, migration 0061). The cascade, and the bus holder added
to a named push, skip a machine any of whose modules would move to a
build its policy records or an open plan has not sent it, and say which
module, which build, why, and that `push <node>` sends it. A machine
whose last send was not recorded is held until it is named. The named
machine itself, a whole-mesh push and `push --behind` are unchanged.
2026-10-05 22:22:03 +02:00
jochen 9873b3bf13 Keep a replay from moving the mesh backwards, and tighten the queue's edges (review of hq ADR 0219)
A registered replay asked now outranked newer asks of its module, and a plan
took any later outcome as its answer. A replay is now refused while the module
is asked anywhere, a plan module asked under an id is answered by that id
alone, and a rebuild of a commit asks what the module follows. An ask handed
back after a restart no longer reads as dead; a cancel that meets a start is
withdrawn; pause holds for an ask fetched as it lands; kill removes containers
before and after the build ends and says whether its outcome went out; a
holder may say only its own machine is paused.
2026-10-05 19:45:45 +02:00
jochen 541603c15c Announce the build agent's verbs, let a waiting build hear its cancel, and retry a stopped rollout (hq ADR 0219)
The holder's verbs were served and found by nothing; it now answers discovery
with its machine's four, as a runtime announces a seat's verb, so the console
finds <node>/node-build-agent.kill. A cancel publishes no outcome, so a build
waited for also looks at the cancelled set. A plan that stopped at its first
machine is retried by sending that machine the module again, unless a newer
plan holds it.
2026-10-05 19:26:40 +02:00
jochen e610f2d92c Let a build agent be paused, have a build killed, and end an ask cancelled as it took it (hq ADR 0219)
A queued ask could only be waited out and a running build only ended by
stopping the machine, which redelivered it elsewhere. The holder now serves
current, kill, pause and resume on its own machine's subjects; a kill ends
the build's process group and labelled containers and settles the ask as
failed, killed by hand; pause is kept in the workspace across a restart and
said on the bus. The controller writes cancelled ids to a cancelled set the
holder reads on taking an ask, closing the race a delete alone leaves.
2026-10-05 19:17:42 +02:00
jochen 0b07e68cb8 Publish the followed events once the controller's consumer exists
Made from now since issue 248, the consumer does not replay what was
published before it; the test raced the controller's start and published
first.
2026-10-05 18:04:40 +02:00
jschoubben 134d039ff8 Take no dry run in: mark it on the request, echo it on the outcome, set it aside
A dry run of an unreviewed branch was heard by the daemon like any build, registered, and its
definition reached a machine (novox/hq issue 240). The mark now travels with the build and the
daemon records, registers and plans nothing for it.
2026-10-05 09:50:20 +02:00
jochen e11caecdad Let two controllers overlap safely while one hands over to the other (hq issue 213)
The controller's machine moves it from the container to a process by
starting the process first and removing the container once the process
is up (mesh-host's `replaces`). For that moment two controllers share the
store and the bus. Checked what each does:

- the seat's verbs: a queue group per seat, each call answered once. Safe.
- the controller's consumers on CONTROL and EVENTS: push consumers with
  no delivery group, so the second bind is refused with "consumer is
  already bound" and serve exited. The process would restart for ever,
  the host would never see it up, and the container would never go. The
  second controller now stands by and binds when the first lets go
  (tested on a real bus; fails without the change).
- plans: read, changed and saved whole by the 30s timer, by build
  outcomes, by a merge and by `plans stop`. Two timers would each ask a
  tier the other had just asked. Working the plans now takes a
  session-level advisory lock on the inventory: the timer skips while
  another holds it, the other paths wait for it. Build asks happen only
  inside plan work and are covered by the same lock.
2026-10-04 01:11:26 +02:00
jochen 9745c1ab31 An older build request never replaces a newer one's artifact
Builds of one module in flight together finish in any order, and the mesh
took whatever it heard last as what the module is: RegisterModule overwrote
the module's manifest unconditionally, and Held/BuiltAgainst/ReadRepositories
ordered builds by when they were recorded. A postgres build asked before the
mesh-tools runtime fix finished after the one asked after it, and the next
push deployed the stale image (novox/hq issue 219).

A build is now ordered by when it was asked, read from the build-<nanos> id
the controller writes: build.asked and module.built_asked (migration 0055).
A registration from an earlier request than the module's current one is
recorded and refused as superseded. A plan takes as its outcome only a build
asked at or after its own ask, so an earlier plan's leftover build cannot
settle a later plan. Ids of any other shape keep the old order.
2026-10-04 00:22:11 +02:00
jochen 58b4fcb8c8 A bundle stands on the toolchain it is compiled in (hq issue 211)
A manifest names its toolchain by language, not in build.on, so the planner did not know a bundle
depends on the module that publishes its toolchain and built the two in one tier: the bundle
against the old toolchain, recorded as built from the new commit. The edge is read from the
manifest, so it holds before any build recorded it, and a toolchain that moves rebuilds every
bundle compiled in it.
2026-10-03 22:18:20 +02:00
jochen e67c58cd98 Every serving principal may answer the services discovery for what it serves; the controller announces its seat (hq ADR 0197)
Grants: a principal that serves tools subscribes $SRV.PING/$SRV.INFO and those questions under
each name it serves — its own and no other's; the tool runtime and people may ask. The controller
answers discovery for the mesh-controller seat in NATS's services format, one endpoint per verb it
serves, with the seat's description and schema. module list --json says which modules declare tools,
so the console expects an announcement only from those.
2026-10-03 22:11:00 +02:00
jochen 59166b1031 An idle build machine's empty fetch is asked again, not read as the end (hq ADR 0190)
A fetch on a context without a deadline waits the client's own while and reports the deadline
passed — the client's, not ours — and the loop read it as "stop": every idle build agent exited
clean every half minute and was restarted by its supervisor, a crash loop with nothing in the log
to say why. Only our own context ending ends the machine; an empty fetch, however it is reported,
is asked again.
2026-10-03 11:41:23 +02:00
jochen ff5ef0ab60 The controller asks the build role that has a holder, and hears both roles' outcomes (hq ADR 0190, the handover)
A controller that asked node-build-agent from its first run would queue every build where nothing
pulls, and the build that registers build-agent — the first holder — would be among them. So the
role is chosen at ask time from the catalogue: the current role when any assigned module claims it,
the retired one while only the builder does, the current one when neither. Outcomes are followed on
both seats, the controller may publish to both, and a build's log is read under whichever role did
it; a machine on the retired role is proven on the bus to take that role's asks. The switch order
is written where the role is named, and the retired half is marked for removal with the seat row.
2026-10-03 02:51:03 +02:00
jochen a5d6a1187c A build machine serves the seat its credential claims (hq ADR 0190, the handover)
After the build role moved to node-build-agent, nothing would hold it until build-agent is
registered — and registering build-agent needs a build outcome that only the running builder
could produce, bound as it was to the old seat by name. One binary, two roles: the seat a machine
serves is the first its credential claims, as the mesh writes the claims beside the credential it
issues (ADR 0159); the old builder keeps draining mesh-build-machine, a build-agent takes
node-build-agent, and what each says about a build goes out as that seat's events, so an outcome
is heard where the asker of that seat listens. A credential naming no claim serves the current role.
2026-10-03 02:47:49 +02:00
jochen 905f3363c9 Two machines holding the build role share one queue, and neither is handed an ask while busy (hq ADR 0190)
Against a real bus: three asks, two machines; each takes one, the third waits until one is free
and then goes to that one; a machine that stops leaves nothing taken twice. The redelivery of an
ask a dead machine held is the ack wait's, proven by the hand-back test beside this one.

And the order a live mesh switches over in, written where the role is named: queued builds first,
then this controller, then build-agent assigned where machines build, then the builder and the old
seat's stream forgotten.
2026-10-02 22:35:39 +02:00
jochen 9f9d9b3b25 A seat's holders pull one ask at a time from one shared worker (hq ADR 0190, issue 186)
The worker a holder bound was a push consumer in a queue group with one ask in flight: right for
one holder, and with two it would still be a queue of one — the server hands a pushed ask to
whichever subscriber it picks, busy or not, and the in-flight cap is per consumer, not per holder.
Now the worker is pulled: every machine holding the seat binds the same durable and fetches one
ask when it has finished the last, so an idle machine is the one that takes the next, the asks in
flight are bounded by the holders working, and nothing is delivered that nobody asked for — which
is also what ended the race issue 186 describes. A holder's grants trade the delivery subject for
MSG.NEXT on the worker; the ack grant and the heartbeat that keeps a long build alive stay.

Proven against a real bus: the build round trip, a backlog taken by a machine that arrives later,
and work handed back by one machine coming round again.
2026-10-02 22:34:44 +02:00
jochen bde4b61b3b The build role is the node-scoped seat node-build-agent, and its work is shared by every holder (hq ADR 0190)
One build machine built everything, in a queue of one, because the seat was mesh-scoped and a
mesh seat has one holder. ADR 0190 makes building a node role: node-build-agent, held on every
machine that builds, with the work asked of the role and taken by whichever holder is idle. The
work subject of a node-scoped seat carries no node — that token is for a seat's tools, asked of
one machine (design 33 §4) — so holders on several machines read one queue; a test now says so.

The retired mesh-build-machine row stays while the builder module's registered manifest claims
it: a claim to a seat the mesh no longer defines is refused, and the machine holding it would be
unresolvable until build-agent replaces it. Removed once no manifest claims it.

The installer's genesis template (in the host's repository) still grants the controller the old
seat's subjects; its test here says so until that template names node-build-agent.
2026-10-02 22:31:55 +02:00
jschoubben b5df244096 The mesh says what filters a converged machine: filters kept per node, shown by node show, named by status, and previewed with their fates (hq ADR 0168)
A host reports every table and chain that refuses traffic with its owner,
and a converged machine's found firewall's state. The controller keeps both
on the node's record (migration 0054), shows them on node show, names every
converged machine something other than the mesh filters in status — text
and JSON, and such a machine is not well — and the converge preview lists
what filters the machine with the fate of each: retired with the front end,
left as the runtime's, left as a ban, or left in force and not the mesh's.
What was invisible for eleven hours (issues 144, 145) is said by name.
2026-10-02 12:00:12 +02:00
jschoubben 73a34cc0a7 A take is a comparison: the preview, its refusals, the strays, and the policy said at build (hq ADR 0163)
The host now reports, for every held thing, the facts a take compares; the controller keeps them, and
take puts them beside what the module declares — the found image and its age against the declared one,
the found networks and who else is on them, ports and mounts, a found file's difference from the
declared content — and refuses a downgrade without --downgrade and a differing file without --replace
<path>. Without --yes the comparison is printed and nothing is taken. node show lists the facts and
the strays the machine reports. build and the daemon's take-in say when a module's policy rolls the
result out at once. The own-secret refusal points at the provider form for a required secret.
2026-10-01 21:25:27 +02:00
jschoubben 11b20b10ff The build machine takes one ask at a time, and says so while it builds
With the worker consumer's default of many deliveries in flight, every ask behind the one being
built was delivered at once, left unacknowledged for the length of the build, redelivered after the
ack wait and dropped after the fifth time: on 2026-10-01 twenty-six of forty-three builds asked in two
minutes were never built and the queue read as empty (hq issue 186). The holder's worker now has one
in flight, and a running build tells the bus it is still working, as the controller's long handlers
do, so a build longer than the ack wait is neither redelivered nor counted out.
2026-10-01 16:16:13 +02:00
jschoubben e8e502343f The vault's seat, and a report that carries the machine's profile (hq ADR 0161)
mesh-vault joins the mesh's own set — mesh-scoped, delivering secret — because the controller seals
every minted credential with it, which is the test for a seat of the mesh's own; a second provider is
a second claimant, refused by name (issue 106). A report may carry the machine's profile, detected
again by the apply that reports, and the latest replaces the enrolled one: a machine that switched
its network manager is a machine whose uplink holder lacks a capability at its next push (issue 138).
2026-10-01 15:58:04 +02:00
jschoubben 05b90f966a A refused membership does not stop the controller
A stream publish waits for its acknowledgement as long as its context lives, and the server never
acknowledges a publish it refuses. Issuing memberships after a push used the daemon's own context, so
the one refused membership of 2026-10-01 (hq issue 183) held the controller's receive loop for good:
no report, no build outcome, no merge was heard until a restart (hq issue 185). Issuing one
membership is now bounded to ten seconds, and a push says how many could not be issued and stands —
the machines keep the shape they derive until the next push.
2026-10-01 15:30:04 +02:00
jschoubben 603ad61142 The mesh issues an assignment's subjects: a membership per module on a machine (hq ADR 0160)
For every module on every machine the controller composes what that instance serves — its machine's
address always, the module's plain address in a queue when it is alone or its definition says its
instances are interchangeable — the verbs of the seats it holds at the seats' subjects, where its
events land, and what it may reach, resolved the same way for the modules it invokes. Published
beside the node's declaration on `mesh.assignment.<node>.<module>`, last per subject in a stream
that allows direct reads, and the account may read exactly its own. Composed from the same records
the bus's accounts are, so what a runtime serves and what its account may are one composition.
`instances: interchangeable` is the one fact a definition states for it.

The shape issued is the shape the mesh already had, so nothing moves when the membership arrives;
the runtime that reads it instead of deriving it is the next piece.
2026-10-01 14:43:26 +02:00
jschoubben 076e0ae259 A build is taken in where its outcome is heard, and the build tool answers at once (issue 176)
The console's `build` tool answered "no build machine answered within 0s", handed a forge path to
git as written, and a build heard afterwards was recorded and never registered: recording and
registration lived only in the waiting caller, and the tool did not wait.

Now one function takes a build's outcome in — records it, parses the manifest, refuses a definition
naming an installation, registers the module with its source as the seat and path the request
carried — and both the waiting command and the daemon that follows the role's `built` event call
it. `build --wait 0` asks and returns with the id; `builds --log <id>` follows it. The seat verb
says `--self` for a repository given without a scheme.
2026-10-01 01:27:04 +02:00
jschoubben 17f7cb0d9c A build says what it does on the bus, as it happens (novox/hq ADR 0157)
The build-machine seat emits `started` and `log.<build id>` beside `built`. Every line the builder
speaks — each step, each command with its duration, and on failure the command's own output — goes
to stderr as before and onto the bus under the build's id, one subject per build, kept a week in
EVENTS with every other event. `builds --log <id>` reads it back from the stream with a consumer
that is gone when the reading is done, on the command line and as the controller's seat verb;
`builds` lists each build's id and `build` says the id it asked with.

Lines are core publishes with a sequence number, so a build is not slowed by an ack per line and a
gap is visible; `started` and `built` are awaited into the stream. The seat protocol widens
additively at the controller's next start; the holder's grant follows on the broker node's next
composition.
2026-10-01 00:43:07 +02:00
jschoubben d9dbc9a59a A definition names no installation: the check, an operator's value, a context on the git seat
InstallationProblems judges every value the mesh acts on for a name under a public top-level domain
or a public address, with prose, the world's registries, resolvers and certificate authorities
exempt, and a name a resource means on purpose declared with its reason (names-on-purpose). Run by
module check and a catalogue-wide test, not yet at registration, while the declared list shrinks.
${setting:<key>} fills a file from the assignment's settings and is refused when nothing set it.
A build context may live on the git seat; the request carries the seat's clone base (novox/hq ADR
0112, ADR 0155, issues 122 and 134).
2026-09-30 18:38:06 +02:00
jschoubben 70705ffe45 A holder binds its seat's tools when it may, not only when it starts
The grant is a line in the bus's user list the controller itself composes and a push delivers, so
the first controller to serve its seat started before the list named it and every subscription was
refused for good (2026-09-30). A refused subscription is retried until it holds.
2026-09-30 17:59:23 +02:00
jschoubben e9df5dccab The mesh's own verbs are the mesh-controller seat's tools
A seat's protocol lives in the store (migration 0047; seeded additively), a served verb carries its
description and schema, holding a mesh seat requires serving its verbs, a node-scoped seat's tool
carries the node, and the control plane serves status, nodes, node, modules, seats, builds, plan,
assign, unassign, push, build and tools on its seat by running the same commands (novox/hq ADR 0132,
ADR 0154, design 33). A grant of * reaches a role's tools; seat:<seat>.<verb> grants one.
2026-09-30 17:39:38 +02:00
jschoubben 7683ba8b5b The mesh knows which host runs a machine, and says who is behind another
novox/hq 04-ISSUES/087. A host refuses a declaration carrying a field it
does not know, and refuses it WHOLE — deliberately, because that keeps a
half-understood declaration off a machine. It makes every new declaration
field a flag day: hosts first, then the controller. The mesh had no record
of which host any machine ran, so that order was kept by somebody
remembering it, and a machine that refused for this reason reported a
failure with nothing saying why.

The machine has reported its host version since ADR 0141. The
controller's own copy of the report did not have the field, so it was
unmarshalled into nothing and thrown away on arrival. It has it now,
records it, and shows it in `node show` — "not reported" rather than
blank, because a machine that has not said is not a machine running
nothing.

Status says which machines run an older host than another machine does,
and which is newest. Deliberately disagreement rather than staleness:
nothing delivers a host version yet (ADR 0141, accepted and not built), so
the mesh holds no canonical current version and cannot honestly say a
machine is behind THE host. What it can say is that the oldest host in the
mesh is what the mesh may send.

A machine that has reported nothing is left out rather than called
behind. Versions compare as strings, which suits the timestamps and
commits this mesh uses and is wrong for a scheme where "10" sorts before
"9" — said in the code, at the place that would have to learn.
2026-09-30 08:53:59 +02:00
jschoubben fe5988c536 The filter constrains what arrives from outside, and names no network
The forward chain blocked everything passing through the machine and then allowed
the machine's own containers back by naming their address ranges: 172.16.0.0/12 and
192.168.128.0/17 fixed here, the rest recorded per machine by 0043. Every way of
keeping that list correct fails — a constant describes one machine, a recorded range
goes stale in silence and cannot tell a network the mesh made from one a predecessor
left behind, and generating it from the modules would put half the rule set on the
machine.

The mesh has no position on a container reaching outward: that is not a port opened
to anybody. So both chains are written around the links traffic arrives on. What did
not arrive from outside is accepted in one line; what did meets the declared rules.
The tunnel is named beside the outward links rather than treated as inside, or a port
nothing declares would be reachable from every machine in the mesh.

A machine that has not reported an outward link is sent no filter and keeps the one
it has, refused where a person reads it rather than as a rule set that will not load.

Removes the two constants, `node networks`, and the column behind it. novox/hq ADR
0140, superseding 0137 and 0139.
2026-09-29 01:01:29 +02:00
jschoubben 1c3f44a526 Two faults found on review
A second Accept value would have overwritten the first, because the header was Set per value rather
than Added. One value is all any caller passes today, so nothing was wrong — but a helper that
quietly keeps only the last of what it was given is a trap for whoever passes two.

And a replayed announcement that could not be written was published as an empty body: a fact on the
mesh that says nothing, which the reader can only log and drop. It is now said and skipped, because a
body that cannot be marshalled is this program's fault rather than the bus's.
2026-09-28 16:32:41 +02:00
jschoubben 1ebad3786c The mesh says what it applied, and the replay has an address it may use
The pipeline was observable from a merge to an artifact and went dark where it touched a machine: a
node's report is control traffic only the control plane reads, so nothing said which version a
machine runs, or that it refused to (novox/hq ADR 0134). The control plane now states both under the
seat it holds — a role's events belong to the role and keep their address when the holder is
replaced — and only when the report is news, because a machine reconciles every minute and a fact per
report would be a fact per minute per machine.

Whether a report is news is the store's answer: it holds the previous one, so the listener returns it
and the server states the fact. That also gives the catch-up replay a subject the controller may
publish: it was published as a module's event from a module called "control-plane", which does not
exist, so the controller's own account refused it and every catalogue that asked what it missed was
answered with nothing.
2026-09-28 16:07:18 +02:00
jschoubben 2f3bfda8c0 Work slower than the window says so, and one address is the bus's
Three faults the mesh's own logs showed this morning. A handler that outlives the acknowledgement
window was handed its message again while it was still working: acting on a merge builds modules,
minutes against a thirty-second window, so one merge ran the whole catalogue five times over. The
transport now says the work is in progress while it runs, which is where the window belongs.

Everything the mesh hands out — a token, a membership, a person's credential — took its address
from the enrolment setting, which on a mesh that has moved still names the broker it moved from:
the first person issued after the move was handed the retired broker's port. There is one bus, and
its address is the one the control plane is connected to.

And `operator issue` documented an argument order its parser refused.
2026-09-28 09:51:50 +02:00
jschoubben aa771616bb A merge rebuilds what it changed, and what packages it
Three faults in one path. A merge rebuilt every module built from the repository, so one change in
a repository holding twenty-six of them meant twenty-six builds. A merge into a repository a module
only *packages* source from rebuilt nothing — two modules are built from the control plane's own
repository and neither had ever been rebuilt when it moved — because the manifest the mesh keeps
carries no build section, so a build now says which repositories it read and the mesh keeps that
beside what it stood on. And a module handed over by hand could record a repository with no
directory inside it, which is a module nothing can ever rebuild (novox/hq 04-ISSUES/131, /132).

A change inside no module's own directory is a change to what they share, and everything built from
that repository is rebuilt: rebuilding too much is the safe direction, because the fault this whole
path exists for is a mesh that believes it is current and is not.
2026-09-28 09:20:01 +02:00
jschoubben 0014984116 An older merge does not move a source
The forge announces what it finds merged, and an old merge surfacing late moved the recorded head
backwards and rebuilt everything built from that repository, once per old merge. A merge made
before the source was last seen is history; one that says nothing about when is taken as news.
The catalogue now carries when each source was last seen. The bus-records test follows #116:
a module's tools are every one under its own name.
2026-09-28 05:12:37 +02:00
jschoubben aecac5bda2 One bus: the AMQP transport is gone from the controller
The mesh runs on the seat's bus alone (novox/hq ADR 0131, design 28 task 5.5). The old
transport's consume loop, build request, tool ask, management API and account scoping are
deleted, and the bus switch with them; the controller connects to the broker seat and to
nothing else. The store-window tests keep their assertions on a bus-less fake, and the tests
that only made sense for the old transport's in-memory holding go with it.
2026-09-28 03:36:16 +02:00
jschoubben 81e76fa485 The controller follows the subject it decodes
The decoder named the forge's merge subject as the fourth thing followed and the
list was three long: every message that fell through to that switch panicked the
control plane (2026-09-28). The entry was written and lost between two attempts
at the same edit. A test now walks the list; the composed grants and the genesis
template carry the subject.
2026-09-28 03:13:57 +02:00
jschoubben 525f10b858 A merge on the forge builds what it moved, bases first
The controller follows the forge's merges (novox/hq 04-ISSUES/131). For each
module recorded as built from that repository and branch it records the move to
the merge commit and builds it — bases first, because a module built before the
module it stands on is built against the old one and reports success, and a base
that fails stops what stands on it. Nothing is pushed here: what a finished build
does to the machines running the module stays the upgrade's decision.

Two more things the same ordering gives: `build --behind` builds bases first, and
`build --on <module>` rebuilds everything that stands on a module — the rebuild a
changed base needs, which "behind" does not see because their sources did not
move.
2026-09-28 02:59:02 +02:00
jschoubben 3907ea0db0 The control plane serves and pushes on the bus it is told to
The seams were there and nothing chose a side: serve, push, ask and build all
opened the old bus's connection and declared over its channel, whatever
MESH_BUS_NATS said. So the switch moved every host and left the control plane
unable to follow — "this control plane has no MESH_BROKER_AMQP" with the new bus
named and standing (2026-09-28). That was task 4.3 of design 28, still open.

One place now decides: connectLink reads the switch, refuses both buses named at
once, raises the new bus's streams and this controller's consumers when it is
handed the inventory, and opens the link over whichever bus it is on. Every
caller that sent a declaration or asked a tool through the old channel goes
through the server's bus instead, which the new transport has and the channel is
not. OverNats is that outbound: a declaration is a JetStream publish into the
node's own subject, an event is announced on the subject its name derives to, a
tool is request and reply on the module's tool subject.
2026-09-28 01:29:44 +02:00
jschoubben 497f8ea567 Merge main: the trunk renamed the seats and made them data
Both branches changed the seat set from the same starting point, so every number
collided and every `mesh-*` name existed twice. The trunk's numbers and names win:
this branch's records became 0129/0130 and its migrations 0037/0038, and the
hardcoded rename map gave way to the trunk's `seat_alias` table — a rename is a
row now (ADR 0122), not a recompile.

Three of my checks were wrong and the merge is what showed it:

A seat with an empty protocol is a marker, not an incomplete declaration. Most
node-scoped seats are markers — which module is this machine's packet filter —
and refusing one refused most of the set, the showcase module included. A
mistyped field name is already refused by the parser, so an empty protocol was
written as one deliberately.

A claim on a seat this manifest does not declare is not the parser's to judge. A
module may hold a seat another module declared; that is the whole reason ADR 0126
has callers name the seat and not its provider. Whether the seat exists is a fact
about the catalogue, so the refusal is at registration, where every declaration
is in view.

And a seat may share a name with the provision it delivers. `git`, the npm
registry and the artifact store still do, because renaming a delivering seat
cascades to every consumer requiring it, with a window where a holder stops
resolving mid-flight. The trunk deferred exactly those three on purpose.

Full suite green against a real NATS and store.
2026-09-27 18:50:18 +02:00
jschoubben 5fcde512bc A build is announced under both names on the old bus, or merging breaks the live mesh
Found by asking what merging this would do to the mesh that is actually running — the
only place the question could have been asked, because the tests were green and both
buses were self-consistent.

Moving the build outcome to the role means a catalogue built from the current manifests
listens for the role's name. The catalogue *already running* listens for the module's,
because that is what it was told when it was installed. The two do not meet, so merging
as it stood would have stopped the live mesh's module graph being updated — silently,
since a binding that matches nothing is not an error.

A rename on a live bus needs the publisher and the subscriber to change together, and a
deployment cannot promise which arrives first. So the old bus announces under both names
and the order stops mattering. The module's own name retires with the bus, in step 5's
list; nothing has ever run on the bus being built, so there is no legacy name there and
this doubling has no counterpart.
2026-09-27 17:39:39 +02:00
jschoubben e4e960ec1c A build is work submitted to a role, on both buses
ADR 0121 carried through to working code. `Builders` is the asking side and
`BuildMachine` the taking side, each with an implementation per bus, and the builder
binary and the `build` command now go through them.

On the bus being built, one publish does what two did. The old bus answered the asker
through a reply queue and announced to an events exchange, because two audiences meant
two topologies. Here the outcome is the role's own event: the asker matches it by the
id its request carried, the controller records it, the catalogue places it in the graph.
So a build machine publishes once, needs a reply queue for nothing, and needs a grant
over nobody's inbox — which is what ruled out the alternatives.

The outcome carries the module name now. Only the manifest says what was built, and on
the old bus the separate announcement carried it; with one message for three readers it
belongs in the result. A failed build names none, because it produced no module version
and the catalogue would otherwise place something that was never made.

Checked against a real server: the whole round trip; a third party on the role's event
hearing the same outcome the asker did, which is the claim the decision rests on; work
leaving the queue once settled, so no second machine repeats it; work submitted with no
machine holding the role waiting instead of failing, and being done when one arrives;
and work a machine handed back coming round again.

One thing I got wrong twice now and have written down where it bit: binding to a
consumer must name that consumer's own filter subject, not the narrower subject the
caller cares about. The client compares the two and refuses anything that is not equal,
with "subject does not match consumer".
2026-09-27 16:01:53 +02:00
jschoubben eb72ec36ba 1.7 finished: minting, the file delivered, and a test flake I caused
**First, a correction: the previous commit went in on a false check.** Its message
says the suite passed; it did not. The check piped `go test` through a filter that
swallowed the failures and then printed "green" regardless. Two tests were failing
when 4de10e3 landed.

What was failing was my own doing. Purging the streams instead of deleting them
(4de10e3) left the *consumers* behind, because deleting a stream takes its
consumers with it and purging does not. A durable push consumer surviving between
tests keeps pushing to a delivery subject the previous test's subscription has gone
from: the messages count as delivered, go nowhere, and the next test waits out its
timeout for an announcement the server believes it already sent. Consumers are now
removed with the purge. Five consecutive clean runs.

`-p 1` stays, because two packages asserting and deleting the same fixed-name
objects on one bus is a real race — but its comment said the cause I had guessed
and not the one I found, so it now says the right thing.

**And delivery was not finished when I said it was.** Nothing filled
`Rendering.BusUsers`, so the composed file would never have reached a node.
`composeBusUsers` closes it: composed per push for the machine holding
`mesh-broker`, never kept, because the list is a function of the mesh's records and
a stored copy could disagree with them while both looked consistent. A user with no
credential is left out and named rather than written as a user without a password —
an ordinary situation with an obvious remedy — but a file with no users at all is
refused, because that bus would refuse every connection in the mesh.

**Minting, on both halves.** A node at enrolment and a module at `module issue`.
Three things differ from a management call and each is the point of the move: the
credential is minted into the mesh's records and becomes usable at the next
composition, so no server need be reachable; the password travels beside the address
rather than inside it, because a credential embedded in a URL leaks into every log
line that prints a connection; and a module's durable consumer is derived from what
it declared rather than named, so it cannot ask for delivery of something it did not
say it consumes.

A node reconnecting may be refused until that composition reaches the machine
running the bus. That is what the host's reconnect backoff is for and it is
survivable by design; waiting for the push would hold an enrolment open for as long
as a declaration takes to apply.

Tested that the switch is a switch: a node enrolling on one bus comes away with a
credential for that bus and none for the other, because one that held both could be
half-moved and nothing would say which half.
2026-09-27 03:19:41 +02:00