Commit Graph
338 Commits
Author SHA1 Message Date
jochen 2eb9a22c24 Act under a lease, keep accounts by order, one writer at composition (hq to-be 45 Phase 2)
Two controllers could both act (issue 204), a reconcile's report could
overtake the apply after it and the digest decided (issue 267), and a grant
could make a second writer of a machine's report.

- The lease (internal/lease, ADR 0229): mesh-controller_lease key `holder`,
  15 s age, renewed every 5 s by compare-and-set; the epoch is the revision
  it was taken at. The gate is the clock (stops 3 s before expiry); a refused
  renewal is a loss and the process exits; a holder that stops gives it back.
  serve takes it before asserting the bus. Epochs kept in the store
  (migration 0068 controller_epoch) as a floor: a bucket raised from nothing
  is compacted past it. Unleased (no epoch, S12 urgent) only when nobody
  holds it and the bus will not let it be written. A shell command acts
  under the holder's epoch, or its own lease when none.
- Declarations carry `epoch` inside the signed envelope, only to a machine
  whose latest account carried a report_sequence (mesh-host #35); would-send
  is composed with the epoch last sent. Allot and the send both pass the gate.
- Reports: contract in internal/link/order.go (epoch, sequence,
  report_sequence, older_than, refused_older). Accounts kept by epoch, then
  sequence, then report sequence; older refused, counted; unordered reports
  keep the digest rule. Plans by compare-and-set on a revision, with epoch.
  Conditions and calls carry the epoch and are not written off the lease.
- S12 and S13 (naming the writer by epoch) watched, D5 run; reset of the
  bucket said. Writers table compiled in and enforced in PermissionsFor; the
  controller no longer publishes mesh.control.>. A contract per consumed
  kind, and the empty-on-error lint over the repository.
- mesh-host pinned to its main with the epoch in the validator (D1 validates
  the envelope as sent).

Needs mesh-host's genesis lock with the lease grant (mesh-host PR) for
TestTheInstallersFirstUserListIsWhatTheControllerWouldCompose.
2026-10-06 12:29:18 +02:00
jochen e51c6a2cb9 Replace a value given by hand like one the mesh made (hq ADR 0228)
A given own secret the module reads at start is held by nobody but that
module, so the mesh need not read it to replace it: secret rotate now
works on it, and a value given through secret accept is replaced on its
own after the module's first good start under the mesh. Only a value an
outside party issues (own-secrets "issued-by": "outside") or one the
module applies stays as given, refused with the reason.
2026-10-06 12:13:48 +02:00
jochen b853439792 Assert the bus's objects on every send, not only at start (hq issue 208)
A module or seat holder assigned after the controller started was sent
its declaration and found nothing to bind: messenger on novox
("consumer novox_messenger not found", 2026-10-06) and every first
build-agent holder (2026-10-03). assertBusObjects ran only in the start
raise; a push ensured a module's consumer only when a bus credential was
minted, which a module carried by the runtime never is.

The send's grant now runs the same derivation (assertOnSend) before the
memberships, on every push, cascade, plan send and rotation. A failure
is said in the send's output and raised as bus.objects.unasserted, which
the next send that asserts everything clears; the send itself goes on,
because the objects are the mesh's and holding every machine back for
one would turn one fault into all. Module consumers are each tried and
every failure named. The start raise stays as it was.
2026-10-06 11:45:51 +02:00
jochen 8ddc019cd2 Grant the self-check its ban-list question, say a refusal at once, judge the engine by its delivered version (hq to-be 45 Phase 1)
Live on 2026-10-06, two of the first self-check's findings were its own:

- D8 asked every machine's node-intrusion-prevention.banned, and the
  controller's grant did not name the subject: the bus refused it 24 times
  and D8 timed out after thirty seconds instead of saying so. The verbs the
  self-check asks are named in broker.VerbsTheSelfCheckAsks and granted
  (mesh.seat.<seat>.tool.<verb>.*); each probe declares the seat verbs it
  calls, askSeatTool refuses an undeclared one, and a test over the
  registry fails a probe whose question the controller is not granted.
  AskSeatTool now returns a refused publish at once ("the bus refused…")
  instead of waiting out its timeout; D8 asks the machines in parallel.
- D10 read every machine as behind right after a push: a node-engine says
  its version as the directory it is delivered into, the archive's digest
  (31045596c83a, catalogue versionOf), and D10 compared that with the
  build's commit (1545b00a). It now compares with the versions the
  registered build is delivered as, and a hand-placed engine's commit.
2026-10-06 10:44:38 +02:00
jochen bb1607e424 Say when the mesh is wrong: conditions, watchdogs, the bus's advisories, doctor (hq to-be 45 Phase 1)
Every one of the 48 core failures of research 031 was found by a person
looking; the mesh's answers carried the fact for whoever asked and told
nobody.

- The condition store (to-be 45 §2): mesh-controller_conditions, one key
  per open condition, written by compare-and-set so a person's silence
  and the watchdogs never lose each other's word; every transition kept
  ninety days in mesh-controller_condition-history and said as the
  seat's events condition-raised / condition-changed / condition-cleared
  (the condition at the top level, with event, at, change, why, show),
  offered again while the bus is away. Raised and cleared by observation
  only; a clearing reopened within ten minutes is the same condition with
  its count up, its silence kept. Verbs: conditions, conditions show,
  conditions silence (a hand act, at most a week), conditions history.
- ADR 0224's provider standing is the first kind, provider-failing, held
  by the provider's events; the provider_standing table is no longer read
  or written (left in place: dropping it is the operator's word).
- status leads with the open conditions, urgent first, and says all well
  only with none open; conditions it cannot read are said and not well.
- The signals table compiled in, one watchdog loop over it every 30s: S1
  heartbeat (3 intervals, asleep machines excepted, control node urgent
  after 30 min), S2 report after a send, S3 plan tier, S4 event loop deaf,
  S5 merge not acted, S6 ask lost, S7 call hung, S8 provider silent, S9
  advisories, S10 self-check silent, S11 node tools silent, S13 stale
  refusals; S12, S14, S15 deferred with their reasons. A row that cannot
  see raises probe-failed and clears nothing. A test generated from the
  table suppresses each signal inside and past its bound.
- The bus's advisories (maximum deliveries, a mesh consumer deleted) and
  the controller's own slow consumer and refused subjects, said in the
  mesh's words.
- doctor: the probe registry D1-D10 (D5 deferred) and DW, every five
  minutes, each in thirty seconds; a probe that cannot run is never a
  pass. D1 validates with mesh-host's own validator. Every run ends with
  the doctor-heartbeat event mesh-watcher listens for.
- The controller is granted its new buckets, events, the two advisories
  and $SRV.INFO; the node tools their tools-alive heartbeat. The streams
  and consumers the controller asserts and the ones D6/D7 expect are one
  derivation.
2026-10-06 10:21:11 +02:00
jochen e74c32ed50 Keep calls and hand acts on the bus, answer status at once, record durations (hq to-be 45 Phase 0)
A controller restart lost every call's outcome, `status` composed the mesh
while its caller waited (18.6s live on 2026-10-06, past the 10s window), a
repair by hand left no trace, and the core's bounds had nothing measured to
be set from.

- calls: kept in the controller's bucket mesh-controller_calls (last 1000 or
  14 days, answers bounded to 64 KiB), read by id across a restart; a
  controller starting marks a stopped one's running calls abandoned; each
  call names its caller from the inbox its answer goes to.
- status: the serving controller composes it at start, after news from a
  machine, a build or an acting verb, and every minute; the verb answers the
  last composition at once with when and how long it took. Composing resolves
  each machine once instead of twice.
- hand-act log in mesh-controller_hand-acts: push (required through the seat),
  plans stop/close, broker consumer-reset and the new hand-act record take
  --why/--cause/--condition; `hand-acts` lists them and repeated causes;
  status counts the week's.
- durations (migration 0066): apply (send to first report), heartbeat gap,
  plan tier and build, recorded as heard; `durations` summarises them.
- the controller's seat row takes this binary's definition of its own verbs,
  so the console no longer judges calls against an older build's schema.
- the controller is granted its two buckets' subjects.
2026-10-06 02:59:36 +02:00
mesh-admin 146c48fd96 Merge pull request 'Bound a consumer's identity by the provision it requires (hq issue 263, ADR 0225)' (#76) from fix/263-identity-bound-per-provision into main 2026-10-06 00:28:47 +00:00
jochen 6d620f77c3 Bound a consumer's identity by the provision it requires (hq issue 263)
The one global 20-character bound made every consumer pay an object
store's key length, even for provisions that keep no name, and a single
overflow refused the provider's whole declaration. An offer now states
its own bound (identity: {max, in} or false); unsaid, a provider told its
consumers keeps 20 and one told nothing keeps none. module check judges
every identity on the longest machine name before merge, and a provider
leaves an overflowing consumer out of its grants and composes, with the
consumer named by push, plan and status (ADR 0225).
2026-10-06 02:16:20 +02:00
jochen f8286c063d rotate: narrow a pair credential to one consuming module (hq issue 268)
A machine runs many consumers of one provision, each with its own
credential. When one module leaks its credential, `rotate <provision>
--consumer <machine>` was the narrowest act and replaced every module's
on that machine, restarting all of them. --module (and the verb's
module argument beside provision) rotates only that module's.
2026-10-06 02:13:48 +02:00
mesh-admin d1fc25f682 Merge pull request 'Catch up on merges the bus announced and never handed over (hq issue 266)' (#73) from fix/missed-merges-are-caught-up into main 2026-10-05 23:33:18 +00:00
jochen 59f4d486b1 Catch up on merges the bus announced and never handed over (hq issue 266)
The controller acted only on what its events consumer handed it, so a merge
the bus skipped left modules behind with nothing said. The stream is now read
back every five minutes on a single-filter consumer, and any merge that would
still move a module after ten minutes is said and acted on.
2026-10-06 01:29:21 +02:00
jochen 801552c0eb Answer every seat call within ten seconds and keep what came of it (hq issue 265)
A push outlasted the console's 30s wait and, when it sent the bus its
changed user list, the broker's reload forgot the reply it may send:
the push happened and its caller was told it did not answer. Calls now
answer in full or as running with an id, a push answers before it
sends, refused answers are recorded on their call, and 'calls' reads
them back.
2026-10-06 01:14:58 +02:00
jochen 0f0028785c Refuse a verb argument the seat would pass over, and say a push is of the whole mesh
A push naming one machine reached the verb without it and pushed every
machine behind (hq issue 244). The controller now refuses any argument a
verb does not declare, any it composed its command line without, and a
switch that is not true or false; a push that names no machine says first
that it is the whole mesh. Tests walk every served verb: no argument is
ever ignored, and every flag of a verb's command, read from the source, is
in its schema or accounted for. plan gains files, push behind, builds and
plans limit.
2026-10-06 00:37:14 +02:00
jochen 8d9d33ae85 Report a provider that keeps failing a consumer in status (hq ADR 0224)
The identity provider failed every consumer for a day and status called the
mesh well (hq issue 179). The controller now follows every provider's
provisioner.failing/recovered, keeps the newest failing word per provider,
machine and consumer (migration 0065), and status, its JSON and node show
name it until it recovers. Every module that receives contributions is
granted the two events, so no manifest can forget them.
2026-10-06 00:13:23 +02:00
jochen 843b709b59 The registry-trust test reads the runtime's module, which now writes the trust (ADR 0222) 2026-10-05 22:47:43 +02:00
jochen f506fb34ec Let the mesh's resolver seat have several holders on record
musl takes the first reply from any listed nameserver, so a public fallback
beside the mesh's resolver answered NXDOMAIN for mesh names in every Alpine
container (hq ADR 0223). The fix is two mesh resolvers and no public one, which
needs mesh-dns-resolver held on two machines: a seat can now be replicated,
each holder recorded by 'seat <name> --add', checkClaims accepts every holder
on record and still refuses a second holder of any other mesh seat, a holder
answers its own requirement, and a roster fact gives each replicated seat's
holders, this machine first, so resolv-conf can list them. Migration 0062 keys
a holding by seat and assignment.
2026-10-05 22:42:53 +02:00
mesh-admin c34b937dd3 Merge pull request 'The private network writes nothing into the runtime's file; generated resources are collision-checked (hq ADR 0222, issue 190 — 3 of 3)' (#63) from fix/190-the-overlay-writes-no-runtime-file into main 2026-10-05 20:42:32 +00:00
mesh-admin d84c9699b3 Merge pull request 'Tell a module where a mesh seat's holder is reached: ${seat:<seat>:reach} (hq ADR 0222, issue 190 — 1 of 3)' (#62) from fix/190-seat-reach into main 2026-10-05 20:31:32 +00:00
jochen 2421b82ad2 Keep a named push from sending builds a policy or a plan holds back
A named push flushed every other machine whose declaration differed from
what it was last sent (hq ADR 0083). Under an upgrade policy of `record`,
or a plan still waiting on its first machine (ADR 0218), every machine
running the module differs, so `push <one>` sent the held build to all of
them (hq issue 259).

Each send now records which build of each module it carried
(node.sent_builds, migration 0061). The cascade, and the bus holder added
to a named push, skip a machine any of whose modules would move to a
build its policy records or an open plan has not sent it, and say which
module, which build, why, and that `push <node>` sends it. A machine
whose last send was not recorded is held until it is named. The named
machine itself, a whole-mesh push and `push --behind` are unchanged.
2026-10-05 22:22:03 +02:00
jochen 0ebd48a6a8 The private network writes nothing into the runtime's file (hq issue 190)
daemon.json and docker.service belong to the docker module, which holds node-container-runtime
and now states the registry itself through ${seat:mesh-artifact-store:reach} (hq ADR 0222). The
overlay stops generating registry-trust and registry-trust-reload. A generated resource is now
held to the collision check every module is, so a second writer cannot come back through
computed code; resolution never saw what a generator declares.
2026-10-05 22:19:43 +02:00
jochen 67e291c02a Tell a module where a mesh seat's holder is reached (hq ADR 0222)
The container runtime's module must state the mesh's registry to the runtime it owns, so the
controller can stop writing that into the runtime's file (hq issue 190). ${seat:<seat>:reach}
answers host:port without a binding: nothing required, granted or minted, and the address is
one the mesh already composes into every reference it built. Only mesh-artifact-store is
answered; another seat is refused by name. Unanswered in a file written into as JSON, the empty
member is dropped, so the runtime is never told to trust "".
2026-10-05 22:16:23 +02:00
jochen 9873b3bf13 Keep a replay from moving the mesh backwards, and tighten the queue's edges (review of hq ADR 0219)
A registered replay asked now outranked newer asks of its module, and a plan
took any later outcome as its answer. A replay is now refused while the module
is asked anywhere, a plan module asked under an id is answered by that id
alone, and a rebuild of a commit asks what the module follows. An ask handed
back after a restart no longer reads as dead; a cancel that meets a start is
withdrawn; pause holds for an ask fetched as it lands; kill removes containers
before and after the build ends and says whether its outcome went out; a
holder may say only its own machine is paused.
2026-10-05 19:45:45 +02:00
jochen 541603c15c Announce the build agent's verbs, let a waiting build hear its cancel, and retry a stopped rollout (hq ADR 0219)
The holder's verbs were served and found by nothing; it now answers discovery
with its machine's four, as a runtime announces a seat's verb, so the console
finds <node>/node-build-agent.kill. A cancel publishes no outcome, so a build
waited for also looks at the cancelled set. A plan that stopped at its first
machine is retried by sending that machine the module again, unless a newer
plan holds it.
2026-10-05 19:26:40 +02:00
jochen 106507b1d3 Show and change the build queue through the controller, and have plans follow it (hq ADR 0219)
Nothing showed what waited for a build machine, and an ask could not be
dropped without leaving the plan that made it waiting for ever. New verbs:
queue, cancel, clear, rebuild, replay, kill, pause, resume, and plans retry.
Every ask a person drops is recorded failed through the same take-in as a
failed build; a plan keeps the id it asked each module under and matches
its outcome by it. replay is a dry run unless registered, and registering
an older commit than one registered since needs --older (hq issue 207).
A plan waiting on a seat paused on every holder says so and is not late;
a failed plan can be retried, and a rebuild joins the plan holding the
module instead of running beside it.
2026-10-05 19:17:56 +02:00
jochen e610f2d92c Let a build agent be paused, have a build killed, and end an ask cancelled as it took it (hq ADR 0219)
A queued ask could only be waited out and a running build only ended by
stopping the machine, which redelivered it elsewhere. The holder now serves
current, kill, pause and resume on its own machine's subjects; a kill ends
the build's process group and labelled containers and settles the ask as
failed, killed by hand; pause is kept in the workspace across a restart and
said on the bus. The controller writes cancelled ids to a cancelled set the
holder reads on taking an ask, closing the race a delete alone leaves.
2026-10-05 19:17:42 +02:00
jochen 5dcf33db45 Say a plan's first send in the machine's own time, as every other line does
It is kept in UTC and was printed so: 16:32 beside log lines saying 18:32.
2026-10-05 18:35:13 +02:00
jochen 0d2b304f08 Judge each machine's report from its own send, so a first machine opens the gate
With one machine first a module is sent twice, and the tier gate asked every
machine for a report after the second send: the first machine's report, made
between the two, read as stale and the plan waited for ever (hq issue 256).
Also round a plan's wait to the second, not the minute, so it is not 0s.
2026-10-05 18:29:59 +02:00
jochen 16c5e78fa8 Stop a rollout whose first machine is silent, and choose one that is heard from (hq issue 249, ADR 0218)
A first machine that does not report within the bound now stops the
module's rollout, naming it. The first machine is the first by name heard
from lately; reports are judged by what the store says was last sent. A
plan that ends says what it built and never sent, and a failed send's
error is kept in the plan's note.
2026-10-05 18:17:52 +02:00
jochen ed90771382 Send the bus's machine before the grants, and only when its user list moved (hq issue 249)
Grants first could hold back the very declaration that lets the controller
issue them. The holder now goes first, then buckets and memberships, then
the rest; a membership that fails holds back only its own machine, and an
announcement whose send stopped at its grants is asked again. Whether the
holder must go first is read from a digest of the user list it was last
sent, not its whole declaration. Migration renumbered to 0058.
2026-10-05 18:17:52 +02:00
jochen 672d1f4ca1 Leave an announced move to the plan rolling the module out (hq issue 249)
The catalogue's upgrade announcement sent every machine one after another
without waiting for any to apply, beside the plan that now sends one
machine first. A module an open plan has not finished sending is the
plan's to roll out.
2026-10-05 18:17:52 +02:00
jochen 208901c6cc Let a newer plan supersede the older open plans of its repository (hq issue 254, ADR 0218)
A merge planned without looking at open plans, so two plans worked the
same modules and a stuck plan stayed open for ever. The newer plan folds
in what older plans of the same repository and branch had not built or
sent, and closes them as superseded. A person can close a stuck plan by
id with `plans close <id>`.
2026-10-05 18:17:52 +02:00
jochen 4ac5cfe3a7 Roll a plan's module out to one machine first (hq issue 249, ADR 0218)
A plan sent every machine running a module at once, ignoring the module's
upgrade policy. Unless the policy says together, the first machine by name
is sent, recorded in the plan, and the rest follow only once its report
after the send says it applied; a failed first machine stops the plan.
2026-10-05 18:17:52 +02:00
jochen 22660dc274 Issue grants before code, and the bus's machine first (hq issue 249)
A module's new state reached every machine before the permissions to use
it, which came only with a later push. Memberships and buckets now go
before declarations, the machine holding mesh-broker goes first when its
user list must change, and a grant that fails sends nothing and is an
error so the rollout is retried.
2026-10-05 18:17:52 +02:00
jochen d7f359c498 Read a new module's directory as its own, not as shared code (hq issue 252)
A merge adding a module the mesh has not registered rebuilt every module
built from the repository. A path under a directory known to hold modules
belongs to that module when its module.json is among the changed files.
2026-10-05 18:17:52 +02:00
jochen 01c5ab2aab Hold every kept archive by a manifest so the store's collector keeps it (hq issue 253)
The store's garbage-collect marks only from manifests, and archives were
published as bare blobs, so the first real collection would delete every
archive the mesh keeps. PublishArchive now puts a deterministic OCI holder
manifest (empty config, one layer) beside each archive; the sweep holds every
kept archive before it lets anything go, which backfills existing bare blobs,
and lets go of an archive holder-first. A forgotten module no longer keeps its
five recent builds (ADR 0189). `collection [--json]` reports kept archives
held/unheld and what may be let go, so the dry run can be lifted on evidence.
2026-10-05 18:13:23 +02:00
jochen 5761737687 Name the issue the event consumer fix answers: 248 2026-10-05 17:15:49 +02:00
jochen 89ec48b9a9 Make the controller's event consumer from now, and give a stuck one a reset
A consumer made with the server's default replays everything a stream that
keeps history holds: the controller's EVENTS consumer, re-made that way,
replayed a week of merges and builds one at a time and held every new one
behind them. FromNow makes it start at the end; broker consumer-reset
re-makes a stuck one from now, refusing a work queue. hq issue 244.
2026-10-05 17:15:25 +02:00
jschoubben 134d039ff8 Take no dry run in: mark it on the request, echo it on the outcome, set it aside
A dry run of an unreviewed branch was heard by the daemon like any build, registered, and its
definition reached a machine (novox/hq issue 240). The mark now travels with the build and the
daemon records, registers and plans nothing for it.
2026-10-05 09:50:20 +02:00
jschoubben d90c6ab93a Merge pull request 'The mesh's one resolver: its seat, a provider's address, zones, and a node's hosts file (hq ADR 0194, 0196, 0199)' (#251) from feat/mesh-dns-resolver into main 2026-10-04 15:44:31 +00:00
jochen 4ed1057df3 Compose a bus user only for a module that can read an account
Every assigned module was composed as a bus user, though only one declaring
a broker secret can ever be issued an account; the rest were named on every
status, plan and push as credentials never minted (137 now), burying the
real gaps. Their durable consumers are now derived from what the runtime
carries, so nothing they hear changes. The composed file is unchanged:
those users had no password and were already left out. Fixes hq issue 195.
2026-10-04 17:32:50 +02:00
jschoubben 1f4c67a01b The mesh's one resolver: its seat, a provider's address, zones, and a node's hosts file (hq ADR 0194, 0196, 0199)
- mesh-dns-resolver: a mesh seat delivering wildcard-resolution, so every node's resolver
  configuration resolves to its one holder; node-dns-resolver kept until nothing claims it.
- ${bound:<provision>:address}: the providing machine's private address, for the one consumer
  that cannot use a name — a machine's resolver configuration.
- zone: a module declares the zone it answers and the listen that answers it; the controller
  settles it per node, refuses duplicates and shadowing, and hands the resolver .Zones to forward.
- node-hosts-file: a node seat whose holder owns /etc/hosts, with entries/add/remove.
The resolver tests follow the catalogue: no runtime dns (containers copy the machine's resolvers),
live-restore held by resolv-conf, resolv.conf naming the resolver by address then a public one.
2026-10-04 17:32:08 +02:00
jochen 6af891e358 A contribution depends on the seat that receives it, and a collision is refused at assign (hq ADR 0210, issue 235)
The environment and shell contributions were written nowhere on a node without their holder;
they now derive a dependency on node-environment, node-login-shell or node-display-server, met
and refused as ADR 0207's are. Two modules declaring one package, path or unit made the node
unresolvable after the assignment was recorded; that is refused first now, because no later
assignment can complete it.
2026-10-04 15:45:35 +02:00
jochen 35314175f2 Refuse an unmet seat dependency the catalogue could meet (hq ADR 0207 §4)
status reported no unmet dependency on any node once systemd, pacman and docker
were assigned to all four (to-be 42), which is the condition ADR 0207 set for the
switch. A dependency no catalogue module could meet stays a report before and
after the switch, as assign already said it: there is no remedy to name.
2026-10-04 13:07:45 +02:00
jochen e2622fd031 An act says the unmet seat dependencies of the node it acted on, not the mesh's (hq ADR 0207)
assign and unassign say only what they changed on their node; push <node>
lists that node's, push to many counts each and points at status. The
once-per-change log is the serving controller's alone: a one-shot command
starts with no memory, so it logged every node on every call.
2026-10-04 12:50:25 +02:00
mesh-admin 3ee32970ef Merge pull request 'Seat dependencies (hq ADR 0207), the graphical session's seats and display provisions (ADR 0208), groups from several modules' (#264) from feat/0207-a-module-depends-on-the-seats-that-apply-its-resources into main 2026-10-04 10:43:11 +00:00
jochen 10f948e970 A module depends on the node seats that apply its resources (hq ADR 0207)
Seed node-package-manager and node-container-runtime. Derive each module's
dependencies from its declared service, package and container resources;
judge them over the node's whole set, exempting the foundation. Refuse at
assign (several modules may go on as one act) and at unassign of the last
holder; report at composition in status, behind one switch.
2026-10-04 12:34:11 +02:00
jschoubben 41b20b2782 A grant secret belongs to whoever provisions, and the sweep skips what it will not address
Issue 225. The mesh seals one credential per consumer beside the provider's
contributions file, and wrote it root-owned. That was right while a module's
own code ran in a container as root; ADR 0198 moved that code under the node's
runtime, as the node's account, and the secret stayed root's. On the control
machine two consumers went unprovisioned for three hours and the only sign
was a line reading 'secret not readable yet', 4330 times.

The same sentence is already written for a module's own secrets a few hundred
lines above — 'a root-owned 0600 file is one that process cannot read'. This
is that rule reaching the other kind of secret the mesh writes for a module.

Issue 226. The sweep met a reference recorded with the store's old address,
read 'I will not address this' as 'the store refuses everything', and
collected none of the 1681 it had found. Two changes: references from build
records are read through Recorded, where the provenance is known — not in
LetGo, which cannot tell one registry host from another and must stay strict
— and a reference the sweep will not address is now ErrNotOurs, skipped,
never a reason to stop. Only the store refusing ends a sweep.

make check: the two failures both fail on main as well — the resolver test
(hq 202/203) and the service-manager test, which reads this machine's own
shell environment.
2026-10-04 12:21:49 +02:00
mesh-admin 912e9f4e85 Merge pull request 'Assert every declared state's bucket on each push (hq ADR 0201)' (#262) from fix/buckets-on-push into main 2026-10-04 09:21:25 +00:00
jochen babd7b2f47 Assert every declared state's bucket on each push, before the memberships that name it (novox/hq ADR 0201)
The raise at start was the only place buckets were asserted, so a module
registered and assigned since had none until the control plane restarted —
found on the first module to declare state.
2026-10-04 11:13:34 +02:00
jochen cfac579392 Module state is hq ADR 0201 after all: the derived-value record moved to 0202 on hq main 2026-10-04 11:02:42 +02:00