Commit Graph
123 Commits
Author SHA1 Message Date
jochen 6c7af5c63f Run a verb from the controller's own running image, and refuse what it cannot run as a handover (hq issue 289)
mesh/merge-gate pass: builds build-agent, mesh-controller, route-proxy → ace, g14, novox, shanks; no bus step; every machine composes with the change as it…
mesh/repo-check pass: its merge-check.sh passed
mesh/delivery delivered
The witness moves a running build aside into a directory the controller's user
cannot enter, then deletes it; a verb exec'd from os.Executable() in that window
failed with permission denied. /proc/self/exe stays valid while the process lives.
A verb that still cannot start, or arrives while the controller stops, is refused
with link.ErrHandingOver and marked retry: handing-over.
2026-10-07 02:45:09 +02:00
jochen 1cc6a2d759 Keep what each machine says of what it runs, raise it, and gate on it (hq ADR 0240, to-be 48 Phase A)
The gate judged a module by what the mesh saw from outside, so a container that
crash-looped after it applied passed it. Each machine's node-engine now states
the health of every long-running resource it runs; the controller keeps the
newest statement per machine, raises module.<module>.<machine>.unhealthy on the
second statement in a row, clears it on the first that does not say it, and the
gate passes a module only when every long-running resource of it is stated
healthy since the send. An engine that states nothing is judged as before.
2026-10-07 02:28:16 +02:00
jochen 75213b9091 Say a walk that waits too long, and a delivery's own stalls (hq ADR 0239)
mesh/merge-gate pass: builds build-agent, mesh-controller, route-proxy → ace, g14, novox, shanks; no bus step; every machine composes with the change as it…
mesh/repo-check fail: its merge-check.sh failed: FAIL
mesh/delivery delivered
mesh/delivery-group group feat/mesh-delivery-waits-said delivered: every member is delivered
A walk mesh-delivery never lets go waited for ever with nothing open: S16
says it at 30 minutes, urgent at 4 hours, naming plans go. Phase B: probe
D14 reads the delivery owner's stalled and raises delivery.<id>.stalled,
and H2 takes the table's transition through its close.
2026-10-07 00:16:32 +02:00
jochen 487aa040de Let a walk wait for its delivery's word, and serve the delivery's owner (hq ADR 0239)
mesh/merge-gate pass: every machine composes with the change as it did without (0 of 4 compose)
mesh/delivery delivered
mesh/delivery-group group feat/mesh-delivery delivered: every member is delivered
While the mesh-delivery seat has a holder on record, a merge that moves no
core module opens its walk and asks nothing until mesh-delivery or a person
says go; nothing of it is registered before its turn, so no other send
carries it. The controller keeps the planner, the gate, sending and the
walk, and gains the verbs the owner asks with: delivery-plan, -order,
-check (a group composed as one future state), deliver, delivery-stop,
delivery-walks; every walk kept is said as plan-moved.
2026-10-07 00:01:42 +02:00
jochen 3d05d74400 Publish only a commit on its module's trunk; post a pull request's change plan (hq ADR 0238)
mesh/merge-gate pass: the change touches no module of the mesh's graph
mesh/delivery delivered
mesh/delivery-group group feat/the-graph-decides-what-is-checked delivered: every member is delivered
One commit, one plan: a commit off the trunk — a pull request's head, a branch built by
hand, a rebuild or replay of one — is for checking. The build seat reads from its clone
which branches hold the commit, and the controller records and never registers a build
whose commit is not on the branch the module follows (the repository's default for a
new one), so nothing off the trunk can be sent.

A pull request's check now carries its change plan, computed by the planner: what a
merge would build in which order, what each machine would receive, and what is not an
ordinary send — the bus step, a module waiting for a person, a provider's consumers.
2026-10-06 22:49:55 +02:00
jochen 58ebe590a5 Ask the planner what a pull request reaches; map a changed file onto modules in one place (hq ADR 0238)
touchedBy is now the only mapping of changed files onto modules — touched, added, and
read by no build — and reachOfMerge the planner's whole answer with the dependency walk.
The merge handler, the plan what-if, the merge gate's width and composition, and a pull
request's check all ask it, so planning and gating cannot disagree. The gate composes
the definitions of the modules a merge would rebuild or add, not every one in the tree,
and the check says the dependents a merge would build after them.
2026-10-06 22:34:57 +02:00
jochen b24bb030ec A changed file touches exactly the modules whose build reads it (hq issue 280, ADR 0237)
The builder reads a module's own directory (the repository for one built from its root)
and a repository its recipe packages, nothing else. A file in no module's directory was
read as shared code and rebuilt everything built from the repository: 103 modules for a
merge-check.sh added at the catalogue's root. It now touches nothing, in the merge
handler, the release planner and the pull request's check alike, and the gate says so.
2026-10-06 22:34:57 +02:00
jochen 14127d4878 Let the module graph decide what a pull request's check runs, in two layers (hq ADR 0237)
Every pull request the forge announces is mapped onto the mesh's module graph by the
merge handler's rule (issue 278): touching a module — or adding one — runs the gate
(mesh/merge-gate), its judge chosen by the graph (the controller judges itself, the
node-engine by its validator); a repository of the mesh that touches none runs only its
own merge-check.sh (mesh/repo-check), a warning when it has none. Nothing is left pending:
a repository outside the mesh touching nothing is told so as a pass.

The gate moves out of the per-repository scripts into the build seat, so a script is the
repository's own tests and declares its toolchain (go or typescript). The controller's
manifest names every verb of its seat again (ADR 0132), held by a test.
2026-10-06 22:34:56 +02:00
jschoubben 792352dfad Read a rebuild of an unchanged source as no move, whatever image digest it made (hq issue 280)
mesh/merge-gate error: the check could not run: a throwaway postgres:17-alpine could not be raised: docker run --label mesh.build=build-1791317509716888018…
mesh/delivery delivered
An image is not byte-reproducible, so ADR 0236's 'same artifacts is no move'
never held for one: a catalogue merge that did not touch the bus rebuilt it,
and every send to the control node waited for a planned bus upgrade.

The builder now records a source fingerprint per build (module tree, context
trees, bases and toolchains by digest). A rebuild with the fingerprint of the
build it repeats is registered with that build's artifacts, handed to modules
standing on it, holds no push, demands no bus step, and a plan sends and
gates nothing for it. Identical artifacts remain a second way to be no move.
2026-10-06 22:09:11 +02:00
mesh-admin 1c26235bc1 Merge pull request 'Phase 5 (3/4): every test on a bus of its own, at the release the mesh runs (hq ADR 0237)' (#97) from feat/a-suite-that-cannot-flake into main
mesh/delivery delivered
2026-10-06 19:19:41 +00:00
jochen be92762969 Give every test a bus of its own, at the release the mesh runs (hq ADR 0237)
The live tests reached one shared bus and assert, read and remove the mesh's own objects by
their fixed names, so packages run in parallel deleted what each other read and the suite
passed only one package at a time; a red suite read as noise. internal/testbus starts a server
per test, linked in at the nats-server release go.mod pins, and a test holds that pin to the
catalogue's bus image and to the facts snapshot's bus when there is one, so the tests never run
a bus the mesh does not. The waiter test read a timing (the most connections held at one look)
and now reads the state it means (the fewest held across the wait). make check runs the packages
in parallel under the race detector, with a timeout.
2026-10-06 21:17:02 +02:00
jochen 9c714f00d6 Judge every pull request against the mesh that runs, before it merges (hq ADR 0237, to-be 45 §9)
Every check the mesh had ran after a merge, on a machine: a manifest the node-engine refused
(236), an identity a real machine's name made too long (263). merge-gate raises the mesh as the
facts snapshot says it is and the mesh with the change, each in a throwaway store through the
controller's own records, composes every machine twice and validates it with the node-engine's
validator, and fails what the change breaks, naming the machine's roles and the module - plus a
manifest the judging controller cannot read, a consumer left out of its grant, a module removed
while a machine runs it, a new module the node-engine would refuse; it warns on a wide rebuild.

The forge's new head of a pull request becomes a check the controller asks of the build seat:
the head and, beside it, the controller the mesh runs, the catalogue, the host and the lab; a
throwaway store and bus of the versions the mesh runs; the repository's merge-check.sh in the
mesh's Go toolchain with no container runtime socket; then mesh-lab's replays. The verdict is
said as checked, an error never a pass, and nothing is recorded or registered.
2026-10-06 21:11:26 +02:00
jochen b1e02aca1b A module's directory is never shared code, held or not (hq issue 278)
The merge of mesh-catalog 7f99fb4a rebuilt 103 modules with the build agent in tier 0, and ADR
0236 recorded it as "a change to the build agent rebuilds most of the catalogue". The agent had
not changed: modules/showcase/index.ts had. showcase is the catalogue's reference module, held
by no machine, and its manifest was not in the merge, so whatTheMergeTouched read the file as
shared code and rebuilt everything built from the repository (88 came out byte-identical). The
agent stood first only because everything is built by it.

Whether a directory is a module is a fact of the repository at the merge commit, so the forge's
announcer now says it: module_dirs, the changed files' directories holding a module.json there,
with module_dirs_said. A changed file inside one is that module's business; only a file in no
such directory is shared. An announcer that does not say keeps the old rule. `plans` what-if
takes the same list as module-dirs.

And a regression for the open question: nothing depends on the build agent except by being
built by it, and built-by never widens a plan, so a change to the agent - manifest or program -
rebuilds the agent alone; what moved beside it is ordered after it.
2026-10-06 19:55:25 +02:00
jochen d7bf1bae83 Name the decision this builds: hq ADR 0236 (0235 is the bus's snapshot) 2026-10-06 18:56:54 +02:00
jochen d9289ef6d4 Gate a release plan's first machine and roll a failed build back there (hq ADR 0236)
A build that reported applied was sent everywhere; one that then did nothing, served
no tools or broke its machine's word reached every machine. Now the first machine is
judged by the component's health (the core's definitions, as doctor probes H-*, or a
module's own) three times over two minutes within ten; a failing gate puts the previous
build back there once, marks the build, and says it as a condition and an event.
Upgrades roll out by default; the bus is a planned step; a module deleted at its
source is not built (the public-acme plan failure).
2026-10-06 18:56:54 +02:00
jochen 52af210e47 Derive data protection from a module's declared data (hq ADR 0233)
A module's data section says what it keeps and how precious it is; the backup holder's lines,
binding stickiness, retirement on unassign and D13's conditions follow from it, so issue 273's
empty replacement is said and an unassigned module's data is remembered, not forgotten.
2026-10-06 16:47:49 +02:00
jochen 3b2adda1c8 Say a mark-only provider's retirement as mark only
A provider that cannot disable keeps the consumer reachable until a person
deletes it (hq ADR 0230); the listing and the approval say so instead of
claiming access was disabled.
2026-10-06 14:41:15 +02:00
jochen 68009b16fe Hear what providers retire, and let a person approve, reject and delete (hq ADR 0230)
A provider now waits for a person before retiring more than three consumers
or half of what it holds, and deletes only when asked. The controller is that
person's way in: it keeps waiting and rejected sets as conditions, answers them
with retire approve|reject, lists and deletes retired consumers through the
provider's own tools on its machine, records each act in the hand-act log, and
probes for anything retired longer than thirty days (D11).
2026-10-06 14:41:15 +02:00
jochen 751e39186c Heal what is known, under a brake, and say every repair (hq to-be 45 Phase 3)
Research 031 counted the repairs people made by hand: a push to unstick a
plan waiting on a report, a controller restarted to make an object again, a
plan closed, a consumer re-made from now. Each was the ordinary path taken
again by someone who noticed. The healer registry makes each a registered
response to one condition kind, with a budget, a settle and its event:

- H1 sent-not-reported: ask the machine's node-engine to report again
  (mesh.node.<n>.ask.report); if it does not report what it was sent, send
  it again, never moving a build a policy or a plan holds back
- H2 stalled: close a plan whose wait is superseded or finished
- H3 holder-silent / consumer-lost: the send's own assertion of the bus's
  objects (issue 208's note)
- H4 consumer-behind: consumer-reset, only for a consumer the stream table
  marks resettable (the controller's own events consumer)
- H5 is the identity provider's own repair (ADR 0224 §5), registered only

Success is the observation clearing the condition, never the healer; a spent
budget hands the condition to the operator, urgent, with what was tried, and
no healer touches it again. Every act is begun in the store before it is made
(migration 0070), kept in the condition's tried as "healer Hn" and said as
the seat event healer-acted; a heal is never a hand act. More than twelve acts
in an hour stop every healer until an hour after the last, said urgently.
Only the lease holder heals.

S15 is live: a cause repaired by hand twice in a fortnight raises
healer-wanted, naming the healer that was not enough where one exists. D6's
far-behind finding has its own kind, consumer-behind. Nodes are granted the
question; the controller's grant gains healer-acted (genesis lock in
mesh-host). `healers` lists the registry, the acts and the brake; status
counts the week's heals.
2026-10-06 14:26:26 +02:00
jochen 2eb9a22c24 Act under a lease, keep accounts by order, one writer at composition (hq to-be 45 Phase 2)
Two controllers could both act (issue 204), a reconcile's report could
overtake the apply after it and the digest decided (issue 267), and a grant
could make a second writer of a machine's report.

- The lease (internal/lease, ADR 0229): mesh-controller_lease key `holder`,
  15 s age, renewed every 5 s by compare-and-set; the epoch is the revision
  it was taken at. The gate is the clock (stops 3 s before expiry); a refused
  renewal is a loss and the process exits; a holder that stops gives it back.
  serve takes it before asserting the bus. Epochs kept in the store
  (migration 0068 controller_epoch) as a floor: a bucket raised from nothing
  is compacted past it. Unleased (no epoch, S12 urgent) only when nobody
  holds it and the bus will not let it be written. A shell command acts
  under the holder's epoch, or its own lease when none.
- Declarations carry `epoch` inside the signed envelope, only to a machine
  whose latest account carried a report_sequence (mesh-host #35); would-send
  is composed with the epoch last sent. Allot and the send both pass the gate.
- Reports: contract in internal/link/order.go (epoch, sequence,
  report_sequence, older_than, refused_older). Accounts kept by epoch, then
  sequence, then report sequence; older refused, counted; unordered reports
  keep the digest rule. Plans by compare-and-set on a revision, with epoch.
  Conditions and calls carry the epoch and are not written off the lease.
- S12 and S13 (naming the writer by epoch) watched, D5 run; reset of the
  bucket said. Writers table compiled in and enforced in PermissionsFor; the
  controller no longer publishes mesh.control.>. A contract per consumed
  kind, and the empty-on-error lint over the repository.
- mesh-host pinned to its main with the epoch in the validator (D1 validates
  the envelope as sent).

Needs mesh-host's genesis lock with the lease grant (mesh-host PR) for
TestTheInstallersFirstUserListIsWhatTheControllerWouldCompose.
2026-10-06 12:29:18 +02:00
jochen e51c6a2cb9 Replace a value given by hand like one the mesh made (hq ADR 0228)
A given own secret the module reads at start is held by nobody but that
module, so the mesh need not read it to replace it: secret rotate now
works on it, and a value given through secret accept is replaced on its
own after the module's first good start under the mesh. Only a value an
outside party issues (own-secrets "issued-by": "outside") or one the
module applies stays as given, refused with the reason.
2026-10-06 12:13:48 +02:00
jochen 8ddc019cd2 Grant the self-check its ban-list question, say a refusal at once, judge the engine by its delivered version (hq to-be 45 Phase 1)
Live on 2026-10-06, two of the first self-check's findings were its own:

- D8 asked every machine's node-intrusion-prevention.banned, and the
  controller's grant did not name the subject: the bus refused it 24 times
  and D8 timed out after thirty seconds instead of saying so. The verbs the
  self-check asks are named in broker.VerbsTheSelfCheckAsks and granted
  (mesh.seat.<seat>.tool.<verb>.*); each probe declares the seat verbs it
  calls, askSeatTool refuses an undeclared one, and a test over the
  registry fails a probe whose question the controller is not granted.
  AskSeatTool now returns a refused publish at once ("the bus refused…")
  instead of waiting out its timeout; D8 asks the machines in parallel.
- D10 read every machine as behind right after a push: a node-engine says
  its version as the directory it is delivered into, the archive's digest
  (31045596c83a, catalogue versionOf), and D10 compared that with the
  build's commit (1545b00a). It now compares with the versions the
  registered build is delivered as, and a hand-placed engine's commit.
2026-10-06 10:44:38 +02:00
jochen bb1607e424 Say when the mesh is wrong: conditions, watchdogs, the bus's advisories, doctor (hq to-be 45 Phase 1)
Every one of the 48 core failures of research 031 was found by a person
looking; the mesh's answers carried the fact for whoever asked and told
nobody.

- The condition store (to-be 45 §2): mesh-controller_conditions, one key
  per open condition, written by compare-and-set so a person's silence
  and the watchdogs never lose each other's word; every transition kept
  ninety days in mesh-controller_condition-history and said as the
  seat's events condition-raised / condition-changed / condition-cleared
  (the condition at the top level, with event, at, change, why, show),
  offered again while the bus is away. Raised and cleared by observation
  only; a clearing reopened within ten minutes is the same condition with
  its count up, its silence kept. Verbs: conditions, conditions show,
  conditions silence (a hand act, at most a week), conditions history.
- ADR 0224's provider standing is the first kind, provider-failing, held
  by the provider's events; the provider_standing table is no longer read
  or written (left in place: dropping it is the operator's word).
- status leads with the open conditions, urgent first, and says all well
  only with none open; conditions it cannot read are said and not well.
- The signals table compiled in, one watchdog loop over it every 30s: S1
  heartbeat (3 intervals, asleep machines excepted, control node urgent
  after 30 min), S2 report after a send, S3 plan tier, S4 event loop deaf,
  S5 merge not acted, S6 ask lost, S7 call hung, S8 provider silent, S9
  advisories, S10 self-check silent, S11 node tools silent, S13 stale
  refusals; S12, S14, S15 deferred with their reasons. A row that cannot
  see raises probe-failed and clears nothing. A test generated from the
  table suppresses each signal inside and past its bound.
- The bus's advisories (maximum deliveries, a mesh consumer deleted) and
  the controller's own slow consumer and refused subjects, said in the
  mesh's words.
- doctor: the probe registry D1-D10 (D5 deferred) and DW, every five
  minutes, each in thirty seconds; a probe that cannot run is never a
  pass. D1 validates with mesh-host's own validator. Every run ends with
  the doctor-heartbeat event mesh-watcher listens for.
- The controller is granted its new buckets, events, the two advisories
  and $SRV.INFO; the node tools their tools-alive heartbeat. The streams
  and consumers the controller asserts and the ones D6/D7 expect are one
  derivation.
2026-10-06 10:21:11 +02:00
jochen e74c32ed50 Keep calls and hand acts on the bus, answer status at once, record durations (hq to-be 45 Phase 0)
A controller restart lost every call's outcome, `status` composed the mesh
while its caller waited (18.6s live on 2026-10-06, past the 10s window), a
repair by hand left no trace, and the core's bounds had nothing measured to
be set from.

- calls: kept in the controller's bucket mesh-controller_calls (last 1000 or
  14 days, answers bounded to 64 KiB), read by id across a restart; a
  controller starting marks a stopped one's running calls abandoned; each
  call names its caller from the inbox its answer goes to.
- status: the serving controller composes it at start, after news from a
  machine, a build or an acting verb, and every minute; the verb answers the
  last composition at once with when and how long it took. Composing resolves
  each machine once instead of twice.
- hand-act log in mesh-controller_hand-acts: push (required through the seat),
  plans stop/close, broker consumer-reset and the new hand-act record take
  --why/--cause/--condition; `hand-acts` lists them and repeated causes;
  status counts the week's.
- durations (migration 0066): apply (send to first report), heartbeat gap,
  plan tier and build, recorded as heard; `durations` summarises them.
- the controller's seat row takes this binary's definition of its own verbs,
  so the console no longer judges calls against an older build's schema.
- the controller is granted its two buckets' subjects.
2026-10-06 02:59:36 +02:00
jochen 09c0c6b367 Keep the account of the sent declaration over an older one (hq issue 267)
The last report stored per node decides whether a release plan moves on,
and it was whichever arrived last. A report about a declaration the mesh
has moved past now records what it says about the machine but leaves the
account of the apply alone, so arrival order cannot undo the newer.
2026-10-06 01:45:35 +02:00
mesh-admin d1fc25f682 Merge pull request 'Catch up on merges the bus announced and never handed over (hq issue 266)' (#73) from fix/missed-merges-are-caught-up into main
mesh/delivery held for a person: merged without a passing check: only a person decides that it goes on
2026-10-05 23:33:18 +00:00
jochen 59f4d486b1 Catch up on merges the bus announced and never handed over (hq issue 266)
The controller acted only on what its events consumer handed it, so a merge
the bus skipped left modules behind with nothing said. The stream is now read
back every five minutes on a single-filter consumer, and any merge that would
still move a module after ten minutes is said and acted on.
2026-10-06 01:29:21 +02:00
jochen 801552c0eb Answer every seat call within ten seconds and keep what came of it (hq issue 265)
A push outlasted the console's 30s wait and, when it sent the bus its
changed user list, the broker's reload forgot the reply it may send:
the push happened and its caller was told it did not answer. Calls now
answer in full or as running with an id, a push answers before it
sends, refused answers are recorded on their call, and 'calls' reads
them back.
2026-10-06 01:14:58 +02:00
jochen 8d9d33ae85 Report a provider that keeps failing a consumer in status (hq ADR 0224)
The identity provider failed every consumer for a day and status called the
mesh well (hq issue 179). The controller now follows every provider's
provisioner.failing/recovered, keeps the newest failing word per provider,
machine and consumer (migration 0065), and status, its JSON and node show
name it until it recovers. Every module that receives contributions is
granted the two events, so no manifest can forget them.
2026-10-06 00:13:23 +02:00
jochen 2421b82ad2 Keep a named push from sending builds a policy or a plan holds back
A named push flushed every other machine whose declaration differed from
what it was last sent (hq ADR 0083). Under an upgrade policy of `record`,
or a plan still waiting on its first machine (ADR 0218), every machine
running the module differs, so `push <one>` sent the held build to all of
them (hq issue 259).

Each send now records which build of each module it carried
(node.sent_builds, migration 0061). The cascade, and the bus holder added
to a named push, skip a machine any of whose modules would move to a
build its policy records or an open plan has not sent it, and say which
module, which build, why, and that `push <node>` sends it. A machine
whose last send was not recorded is held until it is named. The named
machine itself, a whole-mesh push and `push --behind` are unchanged.
2026-10-05 22:22:03 +02:00
jochen 9873b3bf13 Keep a replay from moving the mesh backwards, and tighten the queue's edges (review of hq ADR 0219)
A registered replay asked now outranked newer asks of its module, and a plan
took any later outcome as its answer. A replay is now refused while the module
is asked anywhere, a plan module asked under an id is answered by that id
alone, and a rebuild of a commit asks what the module follows. An ask handed
back after a restart no longer reads as dead; a cancel that meets a start is
withdrawn; pause holds for an ask fetched as it lands; kill removes containers
before and after the build ends and says whether its outcome went out; a
holder may say only its own machine is paused.
2026-10-05 19:45:45 +02:00
jochen 541603c15c Announce the build agent's verbs, let a waiting build hear its cancel, and retry a stopped rollout (hq ADR 0219)
The holder's verbs were served and found by nothing; it now answers discovery
with its machine's four, as a runtime announces a seat's verb, so the console
finds <node>/node-build-agent.kill. A cancel publishes no outcome, so a build
waited for also looks at the cancelled set. A plan that stopped at its first
machine is retried by sending that machine the module again, unless a newer
plan holds it.
2026-10-05 19:26:40 +02:00
jochen e610f2d92c Let a build agent be paused, have a build killed, and end an ask cancelled as it took it (hq ADR 0219)
A queued ask could only be waited out and a running build only ended by
stopping the machine, which redelivered it elsewhere. The holder now serves
current, kill, pause and resume on its own machine's subjects; a kill ends
the build's process group and labelled containers and settles the ask as
failed, killed by hand; pause is kept in the workspace across a restart and
said on the bus. The controller writes cancelled ids to a cancelled set the
holder reads on taking an ask, closing the race a delete alone leaves.
2026-10-05 19:17:42 +02:00
jochen 0b07e68cb8 Publish the followed events once the controller's consumer exists
Made from now since issue 248, the consumer does not replay what was
published before it; the test raced the controller's start and published
first.
2026-10-05 18:04:40 +02:00
jschoubben 134d039ff8 Take no dry run in: mark it on the request, echo it on the outcome, set it aside
A dry run of an unreviewed branch was heard by the daemon like any build, registered, and its
definition reached a machine (novox/hq issue 240). The mark now travels with the build and the
daemon records, registers and plans nothing for it.
2026-10-05 09:50:20 +02:00
jochen e11caecdad Let two controllers overlap safely while one hands over to the other (hq issue 213)
The controller's machine moves it from the container to a process by
starting the process first and removing the container once the process
is up (mesh-host's `replaces`). For that moment two controllers share the
store and the bus. Checked what each does:

- the seat's verbs: a queue group per seat, each call answered once. Safe.
- the controller's consumers on CONTROL and EVENTS: push consumers with
  no delivery group, so the second bind is refused with "consumer is
  already bound" and serve exited. The process would restart for ever,
  the host would never see it up, and the container would never go. The
  second controller now stands by and binds when the first lets go
  (tested on a real bus; fails without the change).
- plans: read, changed and saved whole by the 30s timer, by build
  outcomes, by a merge and by `plans stop`. Two timers would each ask a
  tier the other had just asked. Working the plans now takes a
  session-level advisory lock on the inventory: the timer skips while
  another holds it, the other paths wait for it. Build asks happen only
  inside plan work and are covered by the same lock.
2026-10-04 01:11:26 +02:00
jochen 9745c1ab31 An older build request never replaces a newer one's artifact
Builds of one module in flight together finish in any order, and the mesh
took whatever it heard last as what the module is: RegisterModule overwrote
the module's manifest unconditionally, and Held/BuiltAgainst/ReadRepositories
ordered builds by when they were recorded. A postgres build asked before the
mesh-tools runtime fix finished after the one asked after it, and the next
push deployed the stale image (novox/hq issue 219).

A build is now ordered by when it was asked, read from the build-<nanos> id
the controller writes: build.asked and module.built_asked (migration 0055).
A registration from an earlier request than the module's current one is
recorded and refused as superseded. A plan takes as its outcome only a build
asked at or after its own ask, so an earlier plan's leftover build cannot
settle a later plan. Ids of any other shape keep the old order.
2026-10-04 00:22:11 +02:00
jochen 58b4fcb8c8 A bundle stands on the toolchain it is compiled in (hq issue 211)
A manifest names its toolchain by language, not in build.on, so the planner did not know a bundle
depends on the module that publishes its toolchain and built the two in one tier: the bundle
against the old toolchain, recorded as built from the new commit. The edge is read from the
manifest, so it holds before any build recorded it, and a toolchain that moves rebuilds every
bundle compiled in it.
2026-10-03 22:18:20 +02:00
jochen e67c58cd98 Every serving principal may answer the services discovery for what it serves; the controller announces its seat (hq ADR 0197)
Grants: a principal that serves tools subscribes $SRV.PING/$SRV.INFO and those questions under
each name it serves — its own and no other's; the tool runtime and people may ask. The controller
answers discovery for the mesh-controller seat in NATS's services format, one endpoint per verb it
serves, with the seat's description and schema. module list --json says which modules declare tools,
so the console expects an announcement only from those.
2026-10-03 22:11:00 +02:00
jochen 59166b1031 An idle build machine's empty fetch is asked again, not read as the end (hq ADR 0190)
A fetch on a context without a deadline waits the client's own while and reports the deadline
passed — the client's, not ours — and the loop read it as "stop": every idle build agent exited
clean every half minute and was restarted by its supervisor, a crash loop with nothing in the log
to say why. Only our own context ending ends the machine; an empty fetch, however it is reported,
is asked again.
2026-10-03 11:41:23 +02:00
jochen ff5ef0ab60 The controller asks the build role that has a holder, and hears both roles' outcomes (hq ADR 0190, the handover)
A controller that asked node-build-agent from its first run would queue every build where nothing
pulls, and the build that registers build-agent — the first holder — would be among them. So the
role is chosen at ask time from the catalogue: the current role when any assigned module claims it,
the retired one while only the builder does, the current one when neither. Outcomes are followed on
both seats, the controller may publish to both, and a build's log is read under whichever role did
it; a machine on the retired role is proven on the bus to take that role's asks. The switch order
is written where the role is named, and the retired half is marked for removal with the seat row.
2026-10-03 02:51:03 +02:00
jochen a5d6a1187c A build machine serves the seat its credential claims (hq ADR 0190, the handover)
After the build role moved to node-build-agent, nothing would hold it until build-agent is
registered — and registering build-agent needs a build outcome that only the running builder
could produce, bound as it was to the old seat by name. One binary, two roles: the seat a machine
serves is the first its credential claims, as the mesh writes the claims beside the credential it
issues (ADR 0159); the old builder keeps draining mesh-build-machine, a build-agent takes
node-build-agent, and what each says about a build goes out as that seat's events, so an outcome
is heard where the asker of that seat listens. A credential naming no claim serves the current role.
2026-10-03 02:47:49 +02:00
jochen 905f3363c9 Two machines holding the build role share one queue, and neither is handed an ask while busy (hq ADR 0190)
Against a real bus: three asks, two machines; each takes one, the third waits until one is free
and then goes to that one; a machine that stops leaves nothing taken twice. The redelivery of an
ask a dead machine held is the ack wait's, proven by the hand-back test beside this one.

And the order a live mesh switches over in, written where the role is named: queued builds first,
then this controller, then build-agent assigned where machines build, then the builder and the old
seat's stream forgotten.
2026-10-02 22:35:39 +02:00
jochen 9f9d9b3b25 A seat's holders pull one ask at a time from one shared worker (hq ADR 0190, issue 186)
The worker a holder bound was a push consumer in a queue group with one ask in flight: right for
one holder, and with two it would still be a queue of one — the server hands a pushed ask to
whichever subscriber it picks, busy or not, and the in-flight cap is per consumer, not per holder.
Now the worker is pulled: every machine holding the seat binds the same durable and fetches one
ask when it has finished the last, so an idle machine is the one that takes the next, the asks in
flight are bounded by the holders working, and nothing is delivered that nobody asked for — which
is also what ended the race issue 186 describes. A holder's grants trade the delivery subject for
MSG.NEXT on the worker; the ack grant and the heartbeat that keeps a long build alive stay.

Proven against a real bus: the build round trip, a backlog taken by a machine that arrives later,
and work handed back by one machine coming round again.
2026-10-02 22:34:44 +02:00
jochen bde4b61b3b The build role is the node-scoped seat node-build-agent, and its work is shared by every holder (hq ADR 0190)
One build machine built everything, in a queue of one, because the seat was mesh-scoped and a
mesh seat has one holder. ADR 0190 makes building a node role: node-build-agent, held on every
machine that builds, with the work asked of the role and taken by whichever holder is idle. The
work subject of a node-scoped seat carries no node — that token is for a seat's tools, asked of
one machine (design 33 §4) — so holders on several machines read one queue; a test now says so.

The retired mesh-build-machine row stays while the builder module's registered manifest claims
it: a claim to a seat the mesh no longer defines is refused, and the machine holding it would be
unresolvable until build-agent replaces it. Removed once no manifest claims it.

The installer's genesis template (in the host's repository) still grants the controller the old
seat's subjects; its test here says so until that template names node-build-agent.
2026-10-02 22:31:55 +02:00
jschoubben b5df244096 The mesh says what filters a converged machine: filters kept per node, shown by node show, named by status, and previewed with their fates (hq ADR 0168)
A host reports every table and chain that refuses traffic with its owner,
and a converged machine's found firewall's state. The controller keeps both
on the node's record (migration 0054), shows them on node show, names every
converged machine something other than the mesh filters in status — text
and JSON, and such a machine is not well — and the converge preview lists
what filters the machine with the fate of each: retired with the front end,
left as the runtime's, left as a ban, or left in force and not the mesh's.
What was invisible for eleven hours (issues 144, 145) is said by name.
2026-10-02 12:00:12 +02:00
jschoubben 73a34cc0a7 A take is a comparison: the preview, its refusals, the strays, and the policy said at build (hq ADR 0163)
The host now reports, for every held thing, the facts a take compares; the controller keeps them, and
take puts them beside what the module declares — the found image and its age against the declared one,
the found networks and who else is on them, ports and mounts, a found file's difference from the
declared content — and refuses a downgrade without --downgrade and a differing file without --replace
<path>. Without --yes the comparison is printed and nothing is taken. node show lists the facts and
the strays the machine reports. build and the daemon's take-in say when a module's policy rolls the
result out at once. The own-secret refusal points at the provider form for a required secret.
2026-10-01 21:25:27 +02:00
jschoubben 11b20b10ff The build machine takes one ask at a time, and says so while it builds
With the worker consumer's default of many deliveries in flight, every ask behind the one being
built was delivered at once, left unacknowledged for the length of the build, redelivered after the
ack wait and dropped after the fifth time: on 2026-10-01 twenty-six of forty-three builds asked in two
minutes were never built and the queue read as empty (hq issue 186). The holder's worker now has one
in flight, and a running build tells the bus it is still working, as the controller's long handlers
do, so a build longer than the ack wait is neither redelivered nor counted out.
2026-10-01 16:16:13 +02:00
jschoubben e8e502343f The vault's seat, and a report that carries the machine's profile (hq ADR 0161)
mesh-vault joins the mesh's own set — mesh-scoped, delivering secret — because the controller seals
every minted credential with it, which is the test for a seat of the mesh's own; a second provider is
a second claimant, refused by name (issue 106). A report may carry the machine's profile, detected
again by the apply that reports, and the latest replaces the enrolled one: a machine that switched
its network manager is a machine whose uplink holder lacks a capability at its next push (issue 138).
2026-10-01 15:58:04 +02:00
jschoubben 05b90f966a A refused membership does not stop the controller
A stream publish waits for its acknowledgement as long as its context lives, and the server never
acknowledges a publish it refuses. Issuing memberships after a push used the daemon's own context, so
the one refused membership of 2026-10-01 (hq issue 183) held the controller's receive loop for good:
no report, no build outcome, no merge was heard until a restart (hq issue 185). Issuing one
membership is now bounded to ten seconds, and a push says how many could not be issued and stands —
the machines keep the shape they derive until the next push.
2026-10-01 15:30:04 +02:00