Commit Graph
901 Commits
Author SHA1 Message Date
jochen 843b709b59 The registry-trust test reads the runtime's module, which now writes the trust (ADR 0222) 2026-10-05 22:47:43 +02:00
jochen f506fb34ec Let the mesh's resolver seat have several holders on record
musl takes the first reply from any listed nameserver, so a public fallback
beside the mesh's resolver answered NXDOMAIN for mesh names in every Alpine
container (hq ADR 0223). The fix is two mesh resolvers and no public one, which
needs mesh-dns-resolver held on two machines: a seat can now be replicated,
each holder recorded by 'seat <name> --add', checkClaims accepts every holder
on record and still refuses a second holder of any other mesh seat, a holder
answers its own requirement, and a roster fact gives each replicated seat's
holders, this machine first, so resolv-conf can list them. Migration 0062 keys
a holding by seat and assignment.
2026-10-05 22:42:53 +02:00
mesh-admin c34b937dd3 Merge pull request 'The private network writes nothing into the runtime's file; generated resources are collision-checked (hq ADR 0222, issue 190 — 3 of 3)' (#63) from fix/190-the-overlay-writes-no-runtime-file into main 2026-10-05 20:42:32 +00:00
mesh-admin d84c9699b3 Merge pull request 'Tell a module where a mesh seat's holder is reached: ${seat:<seat>:reach} (hq ADR 0222, issue 190 — 1 of 3)' (#62) from fix/190-seat-reach into main 2026-10-05 20:31:32 +00:00
mesh-admin 002d5e578c Merge pull request 'A named push sends no build a policy or a plan holds back (hq issue 259, ADR 0221)' (#64) from fix/259-a-named-push-sends-no-held-build into main 2026-10-05 20:26:50 +00:00
jochen 2421b82ad2 Keep a named push from sending builds a policy or a plan holds back
A named push flushed every other machine whose declaration differed from
what it was last sent (hq ADR 0083). Under an upgrade policy of `record`,
or a plan still waiting on its first machine (ADR 0218), every machine
running the module differs, so `push <one>` sent the held build to all of
them (hq issue 259).

Each send now records which build of each module it carried
(node.sent_builds, migration 0061). The cascade, and the bus holder added
to a named push, skip a machine any of whose modules would move to a
build its policy records or an open plan has not sent it, and say which
module, which build, why, and that `push <node>` sends it. A machine
whose last send was not recorded is held until it is named. The named
machine itself, a whole-mesh push and `push --behind` are unchanged.
2026-10-05 22:22:03 +02:00
jochen 0ebd48a6a8 The private network writes nothing into the runtime's file (hq issue 190)
daemon.json and docker.service belong to the docker module, which holds node-container-runtime
and now states the registry itself through ${seat:mesh-artifact-store:reach} (hq ADR 0222). The
overlay stops generating registry-trust and registry-trust-reload. A generated resource is now
held to the collision check every module is, so a second writer cannot come back through
computed code; resolution never saw what a generator declares.
2026-10-05 22:19:43 +02:00
jochen 67e291c02a Tell a module where a mesh seat's holder is reached (hq ADR 0222)
The container runtime's module must state the mesh's registry to the runtime it owns, so the
controller can stop writing that into the runtime's file (hq issue 190). ${seat:<seat>:reach}
answers host:port without a binding: nothing required, granted or minted, and the address is
one the mesh already composes into every reference it built. Only mesh-artifact-store is
answered; another seat is refused by name. Unanswered in a file written into as JSON, the empty
member is dropped, so the runtime is never told to trust "".
2026-10-05 22:16:23 +02:00
mesh-admin 8a400d165e Merge pull request 'Delete node-dns-resolver; resolver config needs the uplink (hq ADR 0220)' (#61) from feat/resolver-config-needs-the-uplink into main 2026-10-05 20:11:18 +00:00
jochen 53d3cd7ce9 Delete node-dns-resolver and make resolver config need the uplink
Nothing has claimed node-dns-resolver since the mesh moved to one resolver
(hq ADR 0194); seeding never removes a row, so a migration deletes it.

resolv.conf stays the mesh's only while the network manager is told to keep
off it, which the node-uplink holder does (ADR 0117). A seat's Needs makes
that a dependency checked at assignment by the ADR 0207 mechanism (hq ADR
0220). The two-claimants test keeps its intent with a synthetic module now
that resolved-split-dns leaves the catalogue.
2026-10-05 21:57:04 +02:00
mesh-admin 6648e4a5c8 Merge pull request 'The controller writes no /etc/hosts (hq ADR 0199)' (#60) from fix/the-controller-writes-no-hosts-file into main 2026-10-05 19:23:41 +00:00
jochen 7fa2568ce7 The controller writes no /etc/hosts (hq ADR 0199)
/etc/hosts is the file of the node-hosts-file seat's holder; the controller writes into no file
another seat's holder owns, and asks that holder if it ever needs a line there. The private
network's module stops asking for the node-names fact; every machine already asks the mesh's
one resolver for these names, and the host gives the region back at the next push.
2026-10-05 21:23:25 +02:00
mesh-admin 4b382ceffd Merge pull request 'A mesh seat's holder elsewhere answers before this machine's own provider (hq issue 258)' (#59) from fix/a-mesh-seat-answers-before-the-machines-own into main 2026-10-05 19:14:24 +00:00
jochen ce86d09d22 A mesh seat's holder elsewhere answers before this machine's own provider (issue 258)
A mesh-wide provision a machine could answer itself was bound to the local provider, with the
seat's holder and any pin consulted only for a provider on another machine. With every machine
still running its own resolver, each bound its resolver configuration to itself while the mesh's
one resolver was held and pinned elsewhere.
2026-10-05 21:14:01 +02:00
jochen 853be00ebe The runtime's file is the runtime module's: the resolver test expects docker to write live-restore (issue 190, ADR 0196)
The catalogue moves daemon.json's live-restore and the reload from resolv-conf to the docker
module, so no module writes another software's configuration. The test composes docker beside
the resolver modules and refuses resolv-conf writing the runtime's file.
2026-10-05 21:14:01 +02:00
mesh-admin e0ce2236dd Merge pull request 'The build queue is controlled through the controller and the build seat (hq ADR 0219)' (#57) from feat/the-build-queue-is-controlled into main 2026-10-05 18:22:01 +00:00
jochen 9873b3bf13 Keep a replay from moving the mesh backwards, and tighten the queue's edges (review of hq ADR 0219)
A registered replay asked now outranked newer asks of its module, and a plan
took any later outcome as its answer. A replay is now refused while the module
is asked anywhere, a plan module asked under an id is answered by that id
alone, and a rebuild of a commit asks what the module follows. An ask handed
back after a restart no longer reads as dead; a cancel that meets a start is
withdrawn; pause holds for an ask fetched as it lands; kill removes containers
before and after the build ends and says whether its outcome went out; a
holder may say only its own machine is paused.
2026-10-05 19:45:45 +02:00
jochen 541603c15c Announce the build agent's verbs, let a waiting build hear its cancel, and retry a stopped rollout (hq ADR 0219)
The holder's verbs were served and found by nothing; it now answers discovery
with its machine's four, as a runtime announces a seat's verb, so the console
finds <node>/node-build-agent.kill. A cancel publishes no outcome, so a build
waited for also looks at the cancelled set. A plan that stopped at its first
machine is retried by sending that machine the module again, unless a newer
plan holds it.
2026-10-05 19:26:40 +02:00
jochen 106507b1d3 Show and change the build queue through the controller, and have plans follow it (hq ADR 0219)
Nothing showed what waited for a build machine, and an ask could not be
dropped without leaving the plan that made it waiting for ever. New verbs:
queue, cancel, clear, rebuild, replay, kill, pause, resume, and plans retry.
Every ask a person drops is recorded failed through the same take-in as a
failed build; a plan keeps the id it asked each module under and matches
its outcome by it. replay is a dry run unless registered, and registering
an older commit than one registered since needs --older (hq issue 207).
A plan waiting on a seat paused on every holder says so and is not late;
a failed plan can be retried, and a rebuild joins the plan holding the
module instead of running beside it.
2026-10-05 19:17:56 +02:00
jochen e610f2d92c Let a build agent be paused, have a build killed, and end an ask cancelled as it took it (hq ADR 0219)
A queued ask could only be waited out and a running build only ended by
stopping the machine, which redelivered it elsewhere. The holder now serves
current, kill, pause and resume on its own machine's subjects; a kill ends
the build's process group and labelled containers and settles the ask as
failed, killed by hand; pause is kept in the workspace across a restart and
said on the bus. The controller writes cancelled ids to a cancelled set the
holder reads on taking an ask, closing the race a delete alone leaves.
2026-10-05 19:17:42 +02:00
mesh-admin f80b6cdbd1 Merge pull request 'Say a plan's first send in the machine's own time' (#56) from fix/a-plan-says-local-time into main 2026-10-05 16:35:25 +00:00
jochen 5dcf33db45 Say a plan's first send in the machine's own time, as every other line does
It is kept in UTC and was printed so: 16:32 beside log lines saying 18:32.
2026-10-05 18:35:13 +02:00
mesh-admin 6abce7e956 Merge pull request 'Judge each machine's report from its own send, so a first machine opens the gate (hq issue 256)' (#55) from fix/the-gate-reads-each-machine-from-its-own-send into main 2026-10-05 16:31:55 +00:00
jochen 0d2b304f08 Judge each machine's report from its own send, so a first machine opens the gate
With one machine first a module is sent twice, and the tier gate asked every
machine for a report after the second send: the first machine's report, made
between the two, read as stale and the plan waited for ever (hq issue 256).
Also round a plan's wait to the second, not the minute, so it is not 0s.
2026-10-05 18:29:59 +02:00
mesh-admin bdfbd0254f Merge pull request 'Delivery in order: grants before code, one machine first, a newer plan takes over (ADR 0218; hq issues 249, 252, 254)' (#54) from fix/delivery-in-order into main 2026-10-05 16:21:05 +00:00
jochen 16c5e78fa8 Stop a rollout whose first machine is silent, and choose one that is heard from (hq issue 249, ADR 0218)
A first machine that does not report within the bound now stops the
module's rollout, naming it. The first machine is the first by name heard
from lately; reports are judged by what the store says was last sent. A
plan that ends says what it built and never sent, and a failed send's
error is kept in the plan's note.
2026-10-05 18:17:52 +02:00
jochen ed90771382 Send the bus's machine before the grants, and only when its user list moved (hq issue 249)
Grants first could hold back the very declaration that lets the controller
issue them. The holder now goes first, then buckets and memberships, then
the rest; a membership that fails holds back only its own machine, and an
announcement whose send stopped at its grants is asked again. Whether the
holder must go first is read from a digest of the user list it was last
sent, not its whole declaration. Migration renumbered to 0058.
2026-10-05 18:17:52 +02:00
jochen 672d1f4ca1 Leave an announced move to the plan rolling the module out (hq issue 249)
The catalogue's upgrade announcement sent every machine one after another
without waiting for any to apply, beside the plan that now sends one
machine first. A module an open plan has not finished sending is the
plan's to roll out.
2026-10-05 18:17:52 +02:00
jochen 208901c6cc Let a newer plan supersede the older open plans of its repository (hq issue 254, ADR 0218)
A merge planned without looking at open plans, so two plans worked the
same modules and a stuck plan stayed open for ever. The newer plan folds
in what older plans of the same repository and branch had not built or
sent, and closes them as superseded. A person can close a stuck plan by
id with `plans close <id>`.
2026-10-05 18:17:52 +02:00
jochen 4ac5cfe3a7 Roll a plan's module out to one machine first (hq issue 249, ADR 0218)
A plan sent every machine running a module at once, ignoring the module's
upgrade policy. Unless the policy says together, the first machine by name
is sent, recorded in the plan, and the rest follow only once its report
after the send says it applied; a failed first machine stops the plan.
2026-10-05 18:17:52 +02:00
jochen 22660dc274 Issue grants before code, and the bus's machine first (hq issue 249)
A module's new state reached every machine before the permissions to use
it, which came only with a later push. Memberships and buckets now go
before declarations, the machine holding mesh-broker goes first when its
user list must change, and a grant that fails sends nothing and is an
error so the rollout is retried.
2026-10-05 18:17:52 +02:00
jochen d7f359c498 Read a new module's directory as its own, not as shared code (hq issue 252)
A merge adding a module the mesh has not registered rebuilt every module
built from the repository. A path under a directory known to hold modules
belongs to that module when its module.json is among the changed files.
2026-10-05 18:17:52 +02:00
mesh-admin ed25fd68e2 Merge pull request 'Hold every kept archive by a manifest so the store's collector keeps it (hq issue 253)' (#53) from fix/every-kept-archive-is-held into main 2026-10-05 16:15:15 +00:00
jochen 01c5ab2aab Hold every kept archive by a manifest so the store's collector keeps it (hq issue 253)
The store's garbage-collect marks only from manifests, and archives were
published as bare blobs, so the first real collection would delete every
archive the mesh keeps. PublishArchive now puts a deterministic OCI holder
manifest (empty config, one layer) beside each archive; the sweep holds every
kept archive before it lets anything go, which backfills existing bare blobs,
and lets go of an archive holder-first. A forgotten module no longer keeps its
five recent builds (ADR 0189). `collection [--json]` reports kept archives
held/unheld and what may be let go, so the dry run can be lifted on evidence.
2026-10-05 18:13:23 +02:00
mesh-admin bb3cd6437b Merge pull request 'Two tests that main broke: bus users after issue 195, and the event consumer made from now (issue 248)' (#52) from fix/the-bus-user-test-follows-issue-195 into main 2026-10-05 16:10:01 +00:00
jochen 0b07e68cb8 Publish the followed events once the controller's consumer exists
Made from now since issue 248, the consumer does not replay what was
published before it; the test raced the controller's start and published
first.
2026-10-05 18:04:40 +02:00
jochen 625d02862c Test the bus users as issue 195 made them: an account-reading module is one, another is not
The postgres-backed test still expected a module that declares no broker
secret to be a bus user, and failed on main since #270; it skips without a
database, so the change's own run did not see it.
2026-10-05 18:01:28 +02:00
mesh-admin e8478e208b Merge pull request 'Make the controller's event consumer from now, and give a stuck one a reset (hq issue 248)' (#51) from fix/a-consumer-on-a-history-stream-starts-from-now into main 2026-10-05 15:27:23 +00:00
jochen 5761737687 Name the issue the event consumer fix answers: 248 2026-10-05 17:15:49 +02:00
jochen 89ec48b9a9 Make the controller's event consumer from now, and give a stuck one a reset
A consumer made with the server's default replays everything a stream that
keeps history holds: the controller's EVENTS consumer, re-made that way,
replayed a week of merges and builds one at a time and held every new one
behind them. FromNow makes it start at the end; broker consumer-reset
re-makes a stuck one from now, refusing a work queue. hq issue 244.
2026-10-05 17:15:25 +02:00
jschoubben 869fb6d6bf Merge pull request 'The node-backup seat (hq ADR 0214, to-be 43)' (#49) from feat/node-backup into main 2026-10-05 10:12:34 +00:00
jschoubben d25b69178b The node-backup seat: a module contributes its backup, the mesh fills its directories
ADR 0214 / to-be 43: a node seat whose holder keeps nightly restore points of what every module on
the machine declares. A contribution of kind backup may name its module's own directories, filled
per module when placed. The catalogue check refuses a store provider that contributes no backup;
parsing does not, so the providers already running stay readable.
2026-10-05 11:42:40 +02:00
mesh-admin a7bb1e0b1c Merge pull request 'node-message-bus: the machine's D-Bus is a node seat (hq ADR 0215)' (#48) from feat/0215-node-message-bus into main 2026-10-05 09:34:14 +00:00
jochen 5eaed84271 node-message-bus: the machine's D-Bus is a node seat (hq ADR 0215) 2026-10-05 11:34:03 +02:00
jschoubben 326b1aec14 Merge pull request 'A dry-run build is taken in by nothing (hq issue 240)' (#47) from fix/a-dry-run-build-is-not-taken-in into main 2026-10-05 09:30:34 +00:00
jschoubben 134d039ff8 Take no dry run in: mark it on the request, echo it on the outcome, set it aside
A dry run of an unreviewed branch was heard by the daemon like any build, registered, and its
definition reached a machine (novox/hq issue 240). The mark now travels with the build and the
daemon records, registers and plans nothing for it.
2026-10-05 09:50:20 +02:00
jschoubben d90c6ab93a Merge pull request 'The mesh's one resolver: its seat, a provider's address, zones, and a node's hosts file (hq ADR 0194, 0196, 0199)' (#251) from feat/mesh-dns-resolver into main 2026-10-04 15:44:31 +00:00
mesh-admin 6b7d2ec49b Merge pull request 'Compose a bus user only for a module that can read an account (hq issue 195)' (#270) from fix/195-only-modules-that-read-an-account-are-bus-users into main 2026-10-04 15:34:06 +00:00
jochen 4ed1057df3 Compose a bus user only for a module that can read an account
Every assigned module was composed as a bus user, though only one declaring
a broker secret can ever be issued an account; the rest were named on every
status, plan and push as credentials never minted (137 now), burying the
real gaps. Their durable consumers are now derived from what the runtime
carries, so nothing they hear changes. The composed file is unchanged:
those users had no password and were already left out. Fixes hq issue 195.
2026-10-04 17:32:50 +02:00
jschoubben 1f4c67a01b The mesh's one resolver: its seat, a provider's address, zones, and a node's hosts file (hq ADR 0194, 0196, 0199)
- mesh-dns-resolver: a mesh seat delivering wildcard-resolution, so every node's resolver
  configuration resolves to its one holder; node-dns-resolver kept until nothing claims it.
- ${bound:<provision>:address}: the providing machine's private address, for the one consumer
  that cannot use a name — a machine's resolver configuration.
- zone: a module declares the zone it answers and the listen that answers it; the controller
  settles it per node, refuses duplicates and shadowing, and hands the resolver .Zones to forward.
- node-hosts-file: a node seat whose holder owns /etc/hosts, with entries/add/remove.
The resolver tests follow the catalogue: no runtime dns (containers copy the machine's resolvers),
live-restore held by resolv-conf, resolv.conf naming the resolver by address then a public one.
2026-10-04 17:32:08 +02:00