Commit Graph
170 Commits
Author SHA1 Message Date
mesh-admin 1c26235bc1 Merge pull request 'Phase 5 (3/4): every test on a bus of its own, at the release the mesh runs (hq ADR 0237)' (#97) from feat/a-suite-that-cannot-flake into main
mesh/delivery delivered
2026-10-06 19:19:41 +00:00
jochen be92762969 Give every test a bus of its own, at the release the mesh runs (hq ADR 0237)
The live tests reached one shared bus and assert, read and remove the mesh's own objects by
their fixed names, so packages run in parallel deleted what each other read and the suite
passed only one package at a time; a red suite read as noise. internal/testbus starts a server
per test, linked in at the nats-server release go.mod pins, and a test holds that pin to the
catalogue's bus image and to the facts snapshot's bus when there is one, so the tests never run
a bus the mesh does not. The waiter test read a timing (the most connections held at one look)
and now reads the state it means (the fewest held across the wait). make check runs the packages
in parallel under the race detector, with a timeout.
2026-10-06 21:17:02 +02:00
jochen 068283137b Keep a facts snapshot for merge checks, and say when it goes stale (hq to-be 45 Phase 5, S14)
Every check the mesh had was right about the world it was given and none was given
the mesh's: a real machine's name made an identity too long (263), the node-engine
refused what the catalogue check passed (236). The controller now composes what a
check needs - every machine under a pseudonym of its name's length, its roles,
system, builds, capabilities, assignments, pins, settings and how its declaration
composes; every seat, module and source; the bus, store and node-engine versions it
runs - with no secret, no address and no name, and keeps it in the artifact store
as facts:latest when it moved, or daily. The replaced snapshot's manifest is let go
of, so the nightly collector takes it. S14 raises facts-stale past two days.
2026-10-06 20:31:10 +02:00
jochen b1e02aca1b A module's directory is never shared code, held or not (hq issue 278)
The merge of mesh-catalog 7f99fb4a rebuilt 103 modules with the build agent in tier 0, and ADR
0236 recorded it as "a change to the build agent rebuilds most of the catalogue". The agent had
not changed: modules/showcase/index.ts had. showcase is the catalogue's reference module, held
by no machine, and its manifest was not in the merge, so whatTheMergeTouched read the file as
shared code and rebuilt everything built from the repository (88 came out byte-identical). The
agent stood first only because everything is built by it.

Whether a directory is a module is a fact of the repository at the merge commit, so the forge's
announcer now says it: module_dirs, the changed files' directories holding a module.json there,
with module_dirs_said. A changed file inside one is that module's business; only a file in no
such directory is shared. An announcer that does not say keeps the old rule. `plans` what-if
takes the same list as module-dirs.

And a regression for the open question: nothing depends on the build agent except by being
built by it, and built-by never widens a plan, so a change to the agent - manifest or program -
rebuilds the agent alone; what moved beside it is ordered after it.
2026-10-06 19:55:25 +02:00
jochen 2bfa6ae4a0 No build reaches a machine without a gate; a release plan walks what waits (hq ADR 0236)
A send carries the machine's whole declaration, so at the switch to roll the next send of
anything would have carried the old default's backlog, unjudged, to every machine. A gated
send now carries and judges everything waiting on its machine; every other send is refused
or leaves the machine; a release plan walks what waits one machine at a time, the control
node last, and one that fails holds the next until a person releases it.
2026-10-06 19:14:25 +02:00
jochen 41f7b2c152 Tell a rebuild that changes nothing by its artifacts and its manifest (hq ADR 0236) 2026-10-06 18:56:54 +02:00
jochen d7bf1bae83 Name the decision this builds: hq ADR 0236 (0235 is the bus's snapshot) 2026-10-06 18:56:54 +02:00
jochen d9289ef6d4 Gate a release plan's first machine and roll a failed build back there (hq ADR 0236)
A build that reported applied was sent everywhere; one that then did nothing, served
no tools or broke its machine's word reached every machine. Now the first machine is
judged by the component's health (the core's definitions, as doctor probes H-*, or a
module's own) three times over two minutes within ten; a failing gate puts the previous
build back there once, marks the build, and says it as a condition and an event.
Upgrades roll out by default; the bus is a planned step; a module deleted at its
source is not built (the public-acme plan failure).
2026-10-06 18:56:54 +02:00
mesh-admin 7aa98e64ce Merge pull request 'Grant the bus's own module the snapshot API and nothing else (hq ADR 0235)' (#90) from feat/bus-snapshot into main
mesh/delivery delivered
2026-10-06 16:24:53 +00:00
jochen 48581af35c Grant the bus's own module the snapshot API and nothing else (hq ADR 0235)
The night's backup of the bus takes each stream through JetStream's snapshot
API, run by the nats module under its own account. The module holding
mesh-broker is composed that account: stream names and info, the snapshot
request, its flow-control acks, its own inbox — no write, which the writers
table checks. A bus module declaring anything else to say on the bus is
refused by module check rather than silently granted nothing. The genesis
user list is unchanged: the controller's grants are.
2026-10-06 18:20:57 +02:00
jochen dcee8cb5bf Say a machine waiting for its push as waiting, not uncomposable (hq issue 275)
Between assign and push a module's own secrets are not made yet; D1 composed
without making them and raised an urgent 'nothing can be sent' that the next
push resolved silently. D1 now composes as the push would (Foreseeing): a
secret the push makes gets a stand-in and is named, one the push is refused on
is refused with the push's words. Waiting is said only past 30 minutes, as a
warning. D3 and D13 expect a holder only once its machine was sent it and
reported or had ten minutes to.
2026-10-06 18:05:38 +02:00
jochen abf9125689 Measure each item its own way; never compare a partial size (hq ADR 0233)
A walk over a large library every hour loads the array that protects it. An item now says how it
is measured — a bounded daily walk, a dataset's counters, or its top level only — and a size that
is a lower bound is kept as such and never read as a shrink.
2026-10-06 17:00:22 +02:00
jochen 52af210e47 Derive data protection from a module's declared data (hq ADR 0233)
A module's data section says what it keeps and how precious it is; the backup holder's lines,
binding stickiness, retirement on unassign and D13's conditions follow from it, so issue 273's
empty replacement is said and an unassigned module's data is remembered, not forgotten.
2026-10-06 16:47:49 +02:00
mesh-admin c988d6d7be Merge pull request 'Keep a consumer bound where its data is; only a pin moves it (hq ADR 0232, issue 273)' (#86) from fix/a-stateful-binding-moves-only-by-a-person into main
mesh/delivery held for a person: merged without a passing check: only a person decides that it goes on
2026-10-06 13:22:42 +00:00
jochen 8bfaf1523e Keep a consumer bound where its data is; only a pin moves it (hq ADR 0232, issue 273)
Issue 258's fix let a mesh seat's holder elsewhere answer before this machine's own provider. Right
for the resolver, which any provider answers alike; for the store's seat it re-bound every database
consumer on a machine running its own store to the holder on another, each was given a fresh, empty
database there, and nothing said so for twenty hours.

- An offer says whether it keeps its consumers' data (`keeps-consumer-data`); unsaid, a provider
  that grants each consumer a credential does. For such a provision the seat's holder no longer
  overrules a provider beside the consumer; a pin still does.
- Where each such consumer was sent is recorded (migration 0071). A resolution that would bind it
  elsewhere keeps the recorded provider and says the move; one whose provider is gone is refused,
  never answered by another.
- A push says a kept move and raises it as an urgent condition at once; the self-check's D12 raises
  it every run, with a pinned move not yet sent as a warning and any unasked move as urgent.
2026-10-06 15:19:23 +02:00
jochen 18da8b37e9 Retire the networking bundle and what the control plane stops shipping (hq ADR 0226)
networking required mesh-wireguard and nothing else; machines are assigned the network directly.
module forget refuses a provided module, so a retired one is removed at start once no machine has it.
Guard route-proxy's public account directory against a reissue.
2026-10-06 14:59:28 +02:00
jochen 751e39186c Heal what is known, under a brake, and say every repair (hq to-be 45 Phase 3)
Research 031 counted the repairs people made by hand: a push to unstick a
plan waiting on a report, a controller restarted to make an object again, a
plan closed, a consumer re-made from now. Each was the ordinary path taken
again by someone who noticed. The healer registry makes each a registered
response to one condition kind, with a budget, a settle and its event:

- H1 sent-not-reported: ask the machine's node-engine to report again
  (mesh.node.<n>.ask.report); if it does not report what it was sent, send
  it again, never moving a build a policy or a plan holds back
- H2 stalled: close a plan whose wait is superseded or finished
- H3 holder-silent / consumer-lost: the send's own assertion of the bus's
  objects (issue 208's note)
- H4 consumer-behind: consumer-reset, only for a consumer the stream table
  marks resettable (the controller's own events consumer)
- H5 is the identity provider's own repair (ADR 0224 §5), registered only

Success is the observation clearing the condition, never the healer; a spent
budget hands the condition to the operator, urgent, with what was tried, and
no healer touches it again. Every act is begun in the store before it is made
(migration 0070), kept in the condition's tried as "healer Hn" and said as
the seat event healer-acted; a heal is never a hand act. More than twelve acts
in an hour stop every healer until an hour after the last, said urgently.
Only the lease holder heals.

S15 is live: a cause repaired by hand twice in a fortnight raises
healer-wanted, naming the healer that was not enough where one exists. D6's
far-behind finding has its own kind, consumer-behind. Nodes are granted the
question; the controller's grant gains healer-acted (genesis lock in
mesh-host). `healers` lists the registry, the acts and the brake; status
counts the week's heals.
2026-10-06 14:26:26 +02:00
jochen 2eb9a22c24 Act under a lease, keep accounts by order, one writer at composition (hq to-be 45 Phase 2)
Two controllers could both act (issue 204), a reconcile's report could
overtake the apply after it and the digest decided (issue 267), and a grant
could make a second writer of a machine's report.

- The lease (internal/lease, ADR 0229): mesh-controller_lease key `holder`,
  15 s age, renewed every 5 s by compare-and-set; the epoch is the revision
  it was taken at. The gate is the clock (stops 3 s before expiry); a refused
  renewal is a loss and the process exits; a holder that stops gives it back.
  serve takes it before asserting the bus. Epochs kept in the store
  (migration 0068 controller_epoch) as a floor: a bucket raised from nothing
  is compacted past it. Unleased (no epoch, S12 urgent) only when nobody
  holds it and the bus will not let it be written. A shell command acts
  under the holder's epoch, or its own lease when none.
- Declarations carry `epoch` inside the signed envelope, only to a machine
  whose latest account carried a report_sequence (mesh-host #35); would-send
  is composed with the epoch last sent. Allot and the send both pass the gate.
- Reports: contract in internal/link/order.go (epoch, sequence,
  report_sequence, older_than, refused_older). Accounts kept by epoch, then
  sequence, then report sequence; older refused, counted; unordered reports
  keep the digest rule. Plans by compare-and-set on a revision, with epoch.
  Conditions and calls carry the epoch and are not written off the lease.
- S12 and S13 (naming the writer by epoch) watched, D5 run; reset of the
  bucket said. Writers table compiled in and enforced in PermissionsFor; the
  controller no longer publishes mesh.control.>. A contract per consumed
  kind, and the empty-on-error lint over the repository.
- mesh-host pinned to its main with the epoch in the validator (D1 validates
  the envelope as sent).

Needs mesh-host's genesis lock with the lease grant (mesh-host PR) for
TestTheInstallersFirstUserListIsWhatTheControllerWouldCompose.
2026-10-06 12:29:18 +02:00
jochen e51c6a2cb9 Replace a value given by hand like one the mesh made (hq ADR 0228)
A given own secret the module reads at start is held by nobody but that
module, so the mesh need not read it to replace it: secret rotate now
works on it, and a value given through secret accept is replaced on its
own after the module's first good start under the mesh. Only a value an
outside party issues (own-secrets "issued-by": "outside") or one the
module applies stays as given, refused with the reason.
2026-10-06 12:13:48 +02:00
jochen bb1607e424 Say when the mesh is wrong: conditions, watchdogs, the bus's advisories, doctor (hq to-be 45 Phase 1)
Every one of the 48 core failures of research 031 was found by a person
looking; the mesh's answers carried the fact for whoever asked and told
nobody.

- The condition store (to-be 45 §2): mesh-controller_conditions, one key
  per open condition, written by compare-and-set so a person's silence
  and the watchdogs never lose each other's word; every transition kept
  ninety days in mesh-controller_condition-history and said as the
  seat's events condition-raised / condition-changed / condition-cleared
  (the condition at the top level, with event, at, change, why, show),
  offered again while the bus is away. Raised and cleared by observation
  only; a clearing reopened within ten minutes is the same condition with
  its count up, its silence kept. Verbs: conditions, conditions show,
  conditions silence (a hand act, at most a week), conditions history.
- ADR 0224's provider standing is the first kind, provider-failing, held
  by the provider's events; the provider_standing table is no longer read
  or written (left in place: dropping it is the operator's word).
- status leads with the open conditions, urgent first, and says all well
  only with none open; conditions it cannot read are said and not well.
- The signals table compiled in, one watchdog loop over it every 30s: S1
  heartbeat (3 intervals, asleep machines excepted, control node urgent
  after 30 min), S2 report after a send, S3 plan tier, S4 event loop deaf,
  S5 merge not acted, S6 ask lost, S7 call hung, S8 provider silent, S9
  advisories, S10 self-check silent, S11 node tools silent, S13 stale
  refusals; S12, S14, S15 deferred with their reasons. A row that cannot
  see raises probe-failed and clears nothing. A test generated from the
  table suppresses each signal inside and past its bound.
- The bus's advisories (maximum deliveries, a mesh consumer deleted) and
  the controller's own slow consumer and refused subjects, said in the
  mesh's words.
- doctor: the probe registry D1-D10 (D5 deferred) and DW, every five
  minutes, each in thirty seconds; a probe that cannot run is never a
  pass. D1 validates with mesh-host's own validator. Every run ends with
  the doctor-heartbeat event mesh-watcher listens for.
- The controller is granted its new buckets, events, the two advisories
  and $SRV.INFO; the node tools their tools-alive heartbeat. The streams
  and consumers the controller asserts and the ones D6/D7 expect are one
  derivation.
2026-10-06 10:21:11 +02:00
jochen e74c32ed50 Keep calls and hand acts on the bus, answer status at once, record durations (hq to-be 45 Phase 0)
A controller restart lost every call's outcome, `status` composed the mesh
while its caller waited (18.6s live on 2026-10-06, past the 10s window), a
repair by hand left no trace, and the core's bounds had nothing measured to
be set from.

- calls: kept in the controller's bucket mesh-controller_calls (last 1000 or
  14 days, answers bounded to 64 KiB), read by id across a restart; a
  controller starting marks a stopped one's running calls abandoned; each
  call names its caller from the inbox its answer goes to.
- status: the serving controller composes it at start, after news from a
  machine, a build or an acting verb, and every minute; the verb answers the
  last composition at once with when and how long it took. Composing resolves
  each machine once instead of twice.
- hand-act log in mesh-controller_hand-acts: push (required through the seat),
  plans stop/close, broker consumer-reset and the new hand-act record take
  --why/--cause/--condition; `hand-acts` lists them and repeated causes;
  status counts the week's.
- durations (migration 0066): apply (send to first report), heartbeat gap,
  plan tier and build, recorded as heard; `durations` summarises them.
- the controller's seat row takes this binary's definition of its own verbs,
  so the console no longer judges calls against an older build's schema.
- the controller is granted its two buckets' subjects.
2026-10-06 02:59:36 +02:00
jochen 09c0c6b367 Keep the account of the sent declaration over an older one (hq issue 267)
The last report stored per node decides whether a release plan moves on,
and it was whichever arrived last. A report about a declaration the mesh
has moved past now records what it says about the machine but leaves the
account of the apply alone, so arrival order cannot undo the newer.
2026-10-06 01:45:35 +02:00
jochen 8d9d33ae85 Report a provider that keeps failing a consumer in status (hq ADR 0224)
The identity provider failed every consumer for a day and status called the
mesh well (hq issue 179). The controller now follows every provider's
provisioner.failing/recovered, keeps the newest failing word per provider,
machine and consumer (migration 0065), and status, its JSON and node show
name it until it recovers. Every module that receives contributions is
granted the two events, so no manifest can forget them.
2026-10-06 00:13:23 +02:00
mesh-admin cc25baa563 Merge pull request 'Rename node-hosts-file to node-hostname, and refuse one seat claimed under two names (hq ADR 0223 part 3)' (#69) from hostname-module into main
mesh/delivery held for a person: merged without a passing check: only a person decides that it goes on
2026-10-05 22:11:41 +00:00
jochen 69b99eec68 Number the hostname seat migration 0064: it merges after the resolver-config one 2026-10-06 00:07:57 +02:00
jochen 09bd0eec4f Number the resolver-config migration 0063: it merges first 2026-10-06 00:07:54 +02:00
jochen ee99a24f77 Rename node-hosts-file to node-hostname, and refuse one seat claimed under two names (hq ADR 0223)
The seat now covers /etc/hostname too. The migration keeps the old name as
an alias so hosts, still assigned while machines move, holds the same seat.
Claims were compared by spelling, so the old and new module would both have
held it on one machine; they are now compared by the seat they resolve to.
2026-10-05 23:43:15 +02:00
jochen f68521da28 Retire node-resolver-config and the seat need it alone used (hq ADR 0223)
The uplink's holder writes /etc/resolv.conf, so the seat that wrote it and
ADR 0220's dependency of it on the uplink have nothing left to say. The
migration deletes the store's row; nothing holds it once resolv-conf is
unassigned everywhere.
2026-10-05 23:40:39 +02:00
jochen f506fb34ec Let the mesh's resolver seat have several holders on record
musl takes the first reply from any listed nameserver, so a public fallback
beside the mesh's resolver answered NXDOMAIN for mesh names in every Alpine
container (hq ADR 0223). The fix is two mesh resolvers and no public one, which
needs mesh-dns-resolver held on two machines: a seat can now be replicated,
each holder recorded by 'seat <name> --add', checkClaims accepts every holder
on record and still refuses a second holder of any other mesh seat, a holder
answers its own requirement, and a roster fact gives each replicated seat's
holders, this machine first, so resolv-conf can list them. Migration 0062 keys
a holding by seat and assignment.
2026-10-05 22:42:53 +02:00
jochen 2421b82ad2 Keep a named push from sending builds a policy or a plan holds back
A named push flushed every other machine whose declaration differed from
what it was last sent (hq ADR 0083). Under an upgrade policy of `record`,
or a plan still waiting on its first machine (ADR 0218), every machine
running the module differs, so `push <one>` sent the held build to all of
them (hq issue 259).

Each send now records which build of each module it carried
(node.sent_builds, migration 0061). The cascade, and the bus holder added
to a named push, skip a machine any of whose modules would move to a
build its policy records or an open plan has not sent it, and say which
module, which build, why, and that `push <node>` sends it. A machine
whose last send was not recorded is held until it is named. The named
machine itself, a whole-mesh push and `push --behind` are unchanged.
2026-10-05 22:22:03 +02:00
jochen 53d3cd7ce9 Delete node-dns-resolver and make resolver config need the uplink
Nothing has claimed node-dns-resolver since the mesh moved to one resolver
(hq ADR 0194); seeding never removes a row, so a migration deletes it.

resolv.conf stays the mesh's only while the network manager is told to keep
off it, which the node-uplink holder does (ADR 0117). A seat's Needs makes
that a dependency checked at assignment by the ADR 0207 mechanism (hq ADR
0220). The two-claimants test keeps its intent with a synthetic module now
that resolved-split-dns leaves the catalogue.
2026-10-05 21:57:04 +02:00
jochen 106507b1d3 Show and change the build queue through the controller, and have plans follow it (hq ADR 0219)
Nothing showed what waited for a build machine, and an ask could not be
dropped without leaving the plan that made it waiting for ever. New verbs:
queue, cancel, clear, rebuild, replay, kill, pause, resume, and plans retry.
Every ask a person drops is recorded failed through the same take-in as a
failed build; a plan keeps the id it asked each module under and matches
its outcome by it. replay is a dry run unless registered, and registering
an older commit than one registered since needs --older (hq issue 207).
A plan waiting on a seat paused on every holder says so and is not late;
a failed plan can be retried, and a rebuild joins the plan holding the
module instead of running beside it.
2026-10-05 19:17:56 +02:00
jochen ed90771382 Send the bus's machine before the grants, and only when its user list moved (hq issue 249)
Grants first could hold back the very declaration that lets the controller
issue them. The holder now goes first, then buckets and memberships, then
the rest; a membership that fails holds back only its own machine, and an
announcement whose send stopped at its grants is asked again. Whether the
holder must go first is read from a digest of the user list it was last
sent, not its whole declaration. Migration renumbered to 0058.
2026-10-05 18:17:52 +02:00
jochen 208901c6cc Let a newer plan supersede the older open plans of its repository (hq issue 254, ADR 0218)
A merge planned without looking at open plans, so two plans worked the
same modules and a stuck plan stayed open for ever. The newer plan folds
in what older plans of the same repository and branch had not built or
sent, and closes them as superseded. A person can close a stuck plan by
id with `plans close <id>`.
2026-10-05 18:17:52 +02:00
jochen 4ac5cfe3a7 Roll a plan's module out to one machine first (hq issue 249, ADR 0218)
A plan sent every machine running a module at once, ignoring the module's
upgrade policy. Unless the policy says together, the first machine by name
is sent, recorded in the plan, and the rest follow only once its report
after the send says it applied; a failed first machine stops the plan.
2026-10-05 18:17:52 +02:00
jochen 01c5ab2aab Hold every kept archive by a manifest so the store's collector keeps it (hq issue 253)
The store's garbage-collect marks only from manifests, and archives were
published as bare blobs, so the first real collection would delete every
archive the mesh keeps. PublishArchive now puts a deterministic OCI holder
manifest (empty config, one layer) beside each archive; the sweep holds every
kept archive before it lets anything go, which backfills existing bare blobs,
and lets go of an archive holder-first. A forgotten module no longer keeps its
five recent builds (ADR 0189). `collection [--json]` reports kept archives
held/unheld and what may be let go, so the dry run can be lifted on evidence.
2026-10-05 18:13:23 +02:00
jochen 625d02862c Test the bus users as issue 195 made them: an account-reading module is one, another is not
The postgres-backed test still expected a module that declares no broker
secret to be a bus user, and failed on main since #270; it skips without a
database, so the change's own run did not see it.
2026-10-05 18:01:28 +02:00
jochen 4ed1057df3 Compose a bus user only for a module that can read an account
Every assigned module was composed as a bus user, though only one declaring
a broker secret can ever be issued an account; the rest were named on every
status, plan and push as credentials never minted (137 now), burying the
real gaps. Their durable consumers are now derived from what the runtime
carries, so nothing they hear changes. The composed file is unchanged:
those users had no password and were already left out. Fixes hq issue 195.
2026-10-04 17:32:50 +02:00
jschoubben 41b20b2782 A grant secret belongs to whoever provisions, and the sweep skips what it will not address
Issue 225. The mesh seals one credential per consumer beside the provider's
contributions file, and wrote it root-owned. That was right while a module's
own code ran in a container as root; ADR 0198 moved that code under the node's
runtime, as the node's account, and the secret stayed root's. On the control
machine two consumers went unprovisioned for three hours and the only sign
was a line reading 'secret not readable yet', 4330 times.

The same sentence is already written for a module's own secrets a few hundred
lines above — 'a root-owned 0600 file is one that process cannot read'. This
is that rule reaching the other kind of secret the mesh writes for a module.

Issue 226. The sweep met a reference recorded with the store's old address,
read 'I will not address this' as 'the store refuses everything', and
collected none of the 1681 it had found. Two changes: references from build
records are read through Recorded, where the provenance is known — not in
LetGo, which cannot tell one registry host from another and must stay strict
— and a reference the sweep will not address is now ErrNotOurs, skipped,
never a reason to stop. Only the store refusing ends a sweep.

make check: the two failures both fail on main as well — the resolver test
(hq 202/203) and the service-manager test, which reads this machine's own
shell environment.
2026-10-04 12:21:49 +02:00
jochen cfac579392 Module state is hq ADR 0201 after all: the derived-value record moved to 0202 on hq main 2026-10-04 11:02:42 +02:00
jochen 78915f9f7a Module state is hq ADR 0202: 0201 landed first for a provider's derivations 2026-10-04 03:44:50 +02:00
mesh-admin 17b8f14fe1 Merge pull request 'A module's state on the bus: buckets from the catalogue, grants, membership (hq ADR 0201)' (#257) from feat/module-state-on-the-bus into main
mesh/delivery held for a person: merged without a passing check: only a person decides that it goes on
2026-10-04 01:43:36 +00:00
jschoubben 79993fb498 Rebased onto main: ADR 0188 renumbered to 0201, migration 0055 to 0056
The bundles refactor took ADR 0188 on main, so this work's record is 0201 and
every comment citing it moves with it. Main also took migration 0055 (an
older build never replaces a newer), so the store's collected-artifacts table
is 0056 — a number two migrations share is a schema nobody can trust.

make check passes except TestTheResolverIsToldEveryMachineOnTheNetworkAndToldAgainWhenOneLeaves,
which fails on main too and now for two stacked reasons (hq issues 203 and 202).
2026-10-04 02:45:06 +02:00
jochen aec55b7072 A module's state on the bus: buckets from the catalogue, grants, membership (novox/hq ADR 0201)
A manifest names the state it keeps (state) and reads (reads); the controller
asserts a key-value bucket per name on every raise, grants owners write and
readers read (measured against a running server), issues each assignment its
buckets in the membership, and reports buckets nothing declares without
removing them.
2026-10-04 02:40:49 +02:00
jschoubben c7884f5a72 The store keeps what the records name (hq ADR 0189)
The mesh names what may go from its own build records — a digest it did not
record making is never named, which is what keeps the sweep away from the
images genesis pushed. An artifact stays because a definition the mesh holds
names it, or because it belongs to one of the five most recent successful
builds of its module.

internal/artifacts asks the store to let go of one; internal/inventory
decides and remembers (migration 0055); the sweep runs after a build the mesh
recorded, which is when both the bytes and the keep set moved. Never fatal to
a build.

And the manifest side of while-stopped, refused from the definition alone:
no schedule, run-once, a container the module does not declare, itself.
2026-10-04 02:32:19 +02:00
jochen 02e3482eb5 Refuse the tools-container shape for every module (to-be 38 WP4b's last step)
While some thirty modules still stood in that shape, one already registered so was rebuilt without
complaint. Every module has moved since; the exception would only let one move back.
2026-10-04 02:28:54 +02:00
jochen e11caecdad Let two controllers overlap safely while one hands over to the other (hq issue 213)
The controller's machine moves it from the container to a process by
starting the process first and removing the container once the process
is up (mesh-host's `replaces`). For that moment two controllers share the
store and the bus. Checked what each does:

- the seat's verbs: a queue group per seat, each call answered once. Safe.
- the controller's consumers on CONTROL and EVENTS: push consumers with
  no delivery group, so the second bind is refused with "consumer is
  already bound" and serve exited. The process would restart for ever,
  the host would never see it up, and the container would never go. The
  second controller now stands by and binds when the first lets go
  (tested on a real bus; fails without the change).
- plans: read, changed and saved whole by the 30s timer, by build
  outcomes, by a merge and by `plans stop`. Two timers would each ask a
  tier the other had just asked. Working the plans now takes a
  session-level advisory lock on the inventory: the timer skips while
  another holds it, the other paths wait for it. Build asks happen only
  inside plan work and are covered by the same lock.
2026-10-04 01:11:26 +02:00
jochen 9745c1ab31 An older build request never replaces a newer one's artifact
Builds of one module in flight together finish in any order, and the mesh
took whatever it heard last as what the module is: RegisterModule overwrote
the module's manifest unconditionally, and Held/BuiltAgainst/ReadRepositories
ordered builds by when they were recorded. A postgres build asked before the
mesh-tools runtime fix finished after the one asked after it, and the next
push deployed the stale image (novox/hq issue 219).

A build is now ordered by when it was asked, read from the build-<nanos> id
the controller writes: build.asked and module.built_asked (migration 0055).
A registration from an earlier request than the module's current one is
recorded and refused as superseded. A plan takes as its outcome only a build
asked at or after its own ask, so an earlier plan's leftover build cannot
settle a later plan. Ids of any other shape keep the old order.
2026-10-04 00:22:11 +02:00
mesh-admin c0c3c3fed4 Merge pull request 'A TypeScript bundle is one file per entrypoint, bundled in the toolchain (hq ADR 0193); a toolchain follows the SDK it stands on (hq issue 212)' (#247) from feat/a-typescript-bundle-is-one-file into main
mesh/delivery held for a person: merged without a passing check: only a person decides that it goes on
2026-10-03 21:50:44 +00:00
jochen f873c97db5 Issue only the recorded holder a seat held once for the mesh (novox/hq issue 218)
A module claiming a mesh-scoped seat was granted and issued the seat's subjects on every machine
it runs on, so the store's verbs answered from whichever postgres replied first. Where the mesh
records the seat's holder, only that (node, module) is now issued it; the module's own tools are
untouched everywhere.
2026-10-03 23:30:27 +02:00