Commit Graph
313 Commits
Author SHA1 Message Date
jochen 5dcf33db45 Say a plan's first send in the machine's own time, as every other line does
It is kept in UTC and was printed so: 16:32 beside log lines saying 18:32.
2026-10-05 18:35:13 +02:00
jochen 0d2b304f08 Judge each machine's report from its own send, so a first machine opens the gate
With one machine first a module is sent twice, and the tier gate asked every
machine for a report after the second send: the first machine's report, made
between the two, read as stale and the plan waited for ever (hq issue 256).
Also round a plan's wait to the second, not the minute, so it is not 0s.
2026-10-05 18:29:59 +02:00
jochen 16c5e78fa8 Stop a rollout whose first machine is silent, and choose one that is heard from (hq issue 249, ADR 0218)
A first machine that does not report within the bound now stops the
module's rollout, naming it. The first machine is the first by name heard
from lately; reports are judged by what the store says was last sent. A
plan that ends says what it built and never sent, and a failed send's
error is kept in the plan's note.
2026-10-05 18:17:52 +02:00
jochen ed90771382 Send the bus's machine before the grants, and only when its user list moved (hq issue 249)
Grants first could hold back the very declaration that lets the controller
issue them. The holder now goes first, then buckets and memberships, then
the rest; a membership that fails holds back only its own machine, and an
announcement whose send stopped at its grants is asked again. Whether the
holder must go first is read from a digest of the user list it was last
sent, not its whole declaration. Migration renumbered to 0058.
2026-10-05 18:17:52 +02:00
jochen 672d1f4ca1 Leave an announced move to the plan rolling the module out (hq issue 249)
The catalogue's upgrade announcement sent every machine one after another
without waiting for any to apply, beside the plan that now sends one
machine first. A module an open plan has not finished sending is the
plan's to roll out.
2026-10-05 18:17:52 +02:00
jochen 208901c6cc Let a newer plan supersede the older open plans of its repository (hq issue 254, ADR 0218)
A merge planned without looking at open plans, so two plans worked the
same modules and a stuck plan stayed open for ever. The newer plan folds
in what older plans of the same repository and branch had not built or
sent, and closes them as superseded. A person can close a stuck plan by
id with `plans close <id>`.
2026-10-05 18:17:52 +02:00
jochen 4ac5cfe3a7 Roll a plan's module out to one machine first (hq issue 249, ADR 0218)
A plan sent every machine running a module at once, ignoring the module's
upgrade policy. Unless the policy says together, the first machine by name
is sent, recorded in the plan, and the rest follow only once its report
after the send says it applied; a failed first machine stops the plan.
2026-10-05 18:17:52 +02:00
jochen 22660dc274 Issue grants before code, and the bus's machine first (hq issue 249)
A module's new state reached every machine before the permissions to use
it, which came only with a later push. Memberships and buckets now go
before declarations, the machine holding mesh-broker goes first when its
user list must change, and a grant that fails sends nothing and is an
error so the rollout is retried.
2026-10-05 18:17:52 +02:00
jochen d7f359c498 Read a new module's directory as its own, not as shared code (hq issue 252)
A merge adding a module the mesh has not registered rebuilt every module
built from the repository. A path under a directory known to hold modules
belongs to that module when its module.json is among the changed files.
2026-10-05 18:17:52 +02:00
jochen 01c5ab2aab Hold every kept archive by a manifest so the store's collector keeps it (hq issue 253)
The store's garbage-collect marks only from manifests, and archives were
published as bare blobs, so the first real collection would delete every
archive the mesh keeps. PublishArchive now puts a deterministic OCI holder
manifest (empty config, one layer) beside each archive; the sweep holds every
kept archive before it lets anything go, which backfills existing bare blobs,
and lets go of an archive holder-first. A forgotten module no longer keeps its
five recent builds (ADR 0189). `collection [--json]` reports kept archives
held/unheld and what may be let go, so the dry run can be lifted on evidence.
2026-10-05 18:13:23 +02:00
jochen 5761737687 Name the issue the event consumer fix answers: 248 2026-10-05 17:15:49 +02:00
jochen 89ec48b9a9 Make the controller's event consumer from now, and give a stuck one a reset
A consumer made with the server's default replays everything a stream that
keeps history holds: the controller's EVENTS consumer, re-made that way,
replayed a week of merges and builds one at a time and held every new one
behind them. FromNow makes it start at the end; broker consumer-reset
re-makes a stuck one from now, refusing a work queue. hq issue 244.
2026-10-05 17:15:25 +02:00
jschoubben 134d039ff8 Take no dry run in: mark it on the request, echo it on the outcome, set it aside
A dry run of an unreviewed branch was heard by the daemon like any build, registered, and its
definition reached a machine (novox/hq issue 240). The mark now travels with the build and the
daemon records, registers and plans nothing for it.
2026-10-05 09:50:20 +02:00
jschoubben d90c6ab93a Merge pull request 'The mesh's one resolver: its seat, a provider's address, zones, and a node's hosts file (hq ADR 0194, 0196, 0199)' (#251) from feat/mesh-dns-resolver into main 2026-10-04 15:44:31 +00:00
jochen 4ed1057df3 Compose a bus user only for a module that can read an account
Every assigned module was composed as a bus user, though only one declaring
a broker secret can ever be issued an account; the rest were named on every
status, plan and push as credentials never minted (137 now), burying the
real gaps. Their durable consumers are now derived from what the runtime
carries, so nothing they hear changes. The composed file is unchanged:
those users had no password and were already left out. Fixes hq issue 195.
2026-10-04 17:32:50 +02:00
jschoubben 1f4c67a01b The mesh's one resolver: its seat, a provider's address, zones, and a node's hosts file (hq ADR 0194, 0196, 0199)
- mesh-dns-resolver: a mesh seat delivering wildcard-resolution, so every node's resolver
  configuration resolves to its one holder; node-dns-resolver kept until nothing claims it.
- ${bound:<provision>:address}: the providing machine's private address, for the one consumer
  that cannot use a name — a machine's resolver configuration.
- zone: a module declares the zone it answers and the listen that answers it; the controller
  settles it per node, refuses duplicates and shadowing, and hands the resolver .Zones to forward.
- node-hosts-file: a node seat whose holder owns /etc/hosts, with entries/add/remove.
The resolver tests follow the catalogue: no runtime dns (containers copy the machine's resolvers),
live-restore held by resolv-conf, resolv.conf naming the resolver by address then a public one.
2026-10-04 17:32:08 +02:00
jochen 6af891e358 A contribution depends on the seat that receives it, and a collision is refused at assign (hq ADR 0210, issue 235)
The environment and shell contributions were written nowhere on a node without their holder;
they now derive a dependency on node-environment, node-login-shell or node-display-server, met
and refused as ADR 0207's are. Two modules declaring one package, path or unit made the node
unresolvable after the assignment was recorded; that is refused first now, because no later
assignment can complete it.
2026-10-04 15:45:35 +02:00
jochen 35314175f2 Refuse an unmet seat dependency the catalogue could meet (hq ADR 0207 §4)
status reported no unmet dependency on any node once systemd, pacman and docker
were assigned to all four (to-be 42), which is the condition ADR 0207 set for the
switch. A dependency no catalogue module could meet stays a report before and
after the switch, as assign already said it: there is no remedy to name.
2026-10-04 13:07:45 +02:00
jochen e2622fd031 An act says the unmet seat dependencies of the node it acted on, not the mesh's (hq ADR 0207)
assign and unassign say only what they changed on their node; push <node>
lists that node's, push to many counts each and points at status. The
once-per-change log is the serving controller's alone: a one-shot command
starts with no memory, so it logged every node on every call.
2026-10-04 12:50:25 +02:00
mesh-admin 3ee32970ef Merge pull request 'Seat dependencies (hq ADR 0207), the graphical session's seats and display provisions (ADR 0208), groups from several modules' (#264) from feat/0207-a-module-depends-on-the-seats-that-apply-its-resources into main 2026-10-04 10:43:11 +00:00
jochen 10f948e970 A module depends on the node seats that apply its resources (hq ADR 0207)
Seed node-package-manager and node-container-runtime. Derive each module's
dependencies from its declared service, package and container resources;
judge them over the node's whole set, exempting the foundation. Refuse at
assign (several modules may go on as one act) and at unassign of the last
holder; report at composition in status, behind one switch.
2026-10-04 12:34:11 +02:00
jschoubben 41b20b2782 A grant secret belongs to whoever provisions, and the sweep skips what it will not address
Issue 225. The mesh seals one credential per consumer beside the provider's
contributions file, and wrote it root-owned. That was right while a module's
own code ran in a container as root; ADR 0198 moved that code under the node's
runtime, as the node's account, and the secret stayed root's. On the control
machine two consumers went unprovisioned for three hours and the only sign
was a line reading 'secret not readable yet', 4330 times.

The same sentence is already written for a module's own secrets a few hundred
lines above — 'a root-owned 0600 file is one that process cannot read'. This
is that rule reaching the other kind of secret the mesh writes for a module.

Issue 226. The sweep met a reference recorded with the store's old address,
read 'I will not address this' as 'the store refuses everything', and
collected none of the 1681 it had found. Two changes: references from build
records are read through Recorded, where the provenance is known — not in
LetGo, which cannot tell one registry host from another and must stay strict
— and a reference the sweep will not address is now ErrNotOurs, skipped,
never a reason to stop. Only the store refusing ends a sweep.

make check: the two failures both fail on main as well — the resolver test
(hq 202/203) and the service-manager test, which reads this machine's own
shell environment.
2026-10-04 12:21:49 +02:00
mesh-admin 912e9f4e85 Merge pull request 'Assert every declared state's bucket on each push (hq ADR 0201)' (#262) from fix/buckets-on-push into main 2026-10-04 09:21:25 +00:00
jochen babd7b2f47 Assert every declared state's bucket on each push, before the memberships that name it (novox/hq ADR 0201)
The raise at start was the only place buckets were asserted, so a module
registered and assigned since had none until the control plane restarted —
found on the first module to declare state.
2026-10-04 11:13:34 +02:00
jochen cfac579392 Module state is hq ADR 0201 after all: the derived-value record moved to 0202 on hq main 2026-10-04 11:02:42 +02:00
jochen 78915f9f7a Module state is hq ADR 0202: 0201 landed first for a provider's derivations 2026-10-04 03:44:50 +02:00
mesh-admin 17b8f14fe1 Merge pull request 'A module's state on the bus: buckets from the catalogue, grants, membership (hq ADR 0201)' (#257) from feat/module-state-on-the-bus into main 2026-10-04 01:43:36 +00:00
jschoubben b9ad7a2948 Review before merge: refuse a silent disagreement, and bound the sweep
Three things found reading this back, each of which would have been quiet.

A consumer that keeps several holders of one provision (ADR 0094) gets a
login per holder, and a provider derives from the login — so it would make a
resource per holder while the consumer is told one value for the requirement.
That is issue 124's own failure one case to the side: authenticate, then be
refused on every object. Refused now, naming both ends.

The sweep runs inside somebody's build and was unbounded. At most two hundred
artifacts and sixty seconds, stopping at the first refusal because a store
that refuses one refuses all; the rest is offered again next build.

The citation and migration renumbers are in the commit before this one.
2026-10-04 03:27:32 +02:00
jochen cde22ff627 module check names a read of state its owner does not keep, and says what each module keeps and reads (novox/hq ADR 0201) 2026-10-04 02:50:02 +02:00
jochen aec55b7072 A module's state on the bus: buckets from the catalogue, grants, membership (novox/hq ADR 0201)
A manifest names the state it keeps (state) and reads (reads); the controller
asserts a key-value bucket per name on every raise, grants owners write and
readers read (measured against a running server), issues each assignment its
buckets in the membership, and reports buckets nothing declares without
removing them.
2026-10-04 02:40:49 +02:00
jschoubben c7884f5a72 The store keeps what the records name (hq ADR 0189)
The mesh names what may go from its own build records — a digest it did not
record making is never named, which is what keeps the sweep away from the
images genesis pushed. An artifact stays because a definition the mesh holds
names it, or because it belongs to one of the five most recent successful
builds of its module.

internal/artifacts asks the store to let go of one; internal/inventory
decides and remembers (migration 0055); the sweep runs after a build the mesh
recorded, which is when both the bytes and the keep set moved. Never fatal to
a build.

And the manifest side of while-stopped, refused from the definition alone:
no schedule, run-once, a container the module does not declare, itself.
2026-10-04 02:32:19 +02:00
jochen c23be73d4d Run the controller as a Go bundle the host starts as a process (hq issue 213)
The controller is a Go program and was the one piece of the mesh's own Go
code still shipped and run as an image (novox/hq issue 213; ADR 0188 §1:
a module's own code is bundles; §3: a service bundle is a process).

The manifest now builds one Go bundle, `controller`, and runs it as the
process `mesh-controller` (`./mesh-controller serve`) under an account
the module declares. What the container gave it, replaced:

- host network: a process is on the host's network; nothing it reads
  names a container network
- user 65534: the account `mesh-controller`, which owns its secrets and
  its state directory
- the eight mounts: the env names the host paths the mesh already places
  (the store, broker and bus files under the state directory, the
  broker's certificate under /var/lib/mesh-broker-tls); the `broker`
  mount was read by nothing and is gone with the others
- `container-runtime` is no longer required on its machine

Its preparation is the same binary with `prepare`, as a run-once process,
and the process `replaces` the container `server`: the host keeps the
container answering until the process is running (mesh-host). Needs the
previous commit live in the running controller, and the host's
`replaces` on the controller's machine, before it is registered.

No image is built by the mesh any more. The Dockerfile stays for genesis
and the lab (`make image`, its Go base now pinned in the Makefile).
2026-10-04 01:45:25 +02:00
jochen e11caecdad Let two controllers overlap safely while one hands over to the other (hq issue 213)
The controller's machine moves it from the container to a process by
starting the process first and removing the container once the process
is up (mesh-host's `replaces`). For that moment two controllers share the
store and the bus. Checked what each does:

- the seat's verbs: a queue group per seat, each call answered once. Safe.
- the controller's consumers on CONTROL and EVENTS: push consumers with
  no delivery group, so the second bind is refused with "consumer is
  already bound" and serve exited. The process would restart for ever,
  the host would never see it up, and the container would never go. The
  second controller now stands by and binds when the first lets go
  (tested on a real bus; fails without the change).
- plans: read, changed and saved whole by the 30s timer, by build
  outcomes, by a merge and by `plans stop`. Two timers would each ask a
  tier the other had just asked. Working the plans now takes a
  session-level advisory lock on the inventory: the timer skips while
  another holds it, the other paths wait for it. Build asks happen only
  inside plan work and are covered by the same lock.
2026-10-04 01:11:26 +02:00
jochen 9745c1ab31 An older build request never replaces a newer one's artifact
Builds of one module in flight together finish in any order, and the mesh
took whatever it heard last as what the module is: RegisterModule overwrote
the module's manifest unconditionally, and Held/BuiltAgainst/ReadRepositories
ordered builds by when they were recorded. A postgres build asked before the
mesh-tools runtime fix finished after the one asked after it, and the next
push deployed the stale image (novox/hq issue 219).

A build is now ordered by when it was asked, read from the build-<nanos> id
the controller writes: build.asked and module.built_asked (migration 0055).
A registration from an earlier request than the module's current one is
recorded and refused as superseded. A plan takes as its outcome only a build
asked at or after its own ask, so an earlier plan's leftover build cannot
settle a later plan. Ids of any other shape keep the old order.
2026-10-04 00:22:11 +02:00
jochen 74efe8e2e7 Tests follow the grants and issue 203: a person may ask what answers; the resolver test mints its credential 2026-10-03 23:15:57 +02:00
jochen 50cf253a43 Merge remote-tracking branch 'origin/fix/issue-215-a-commit-is-never-a-branch-to-follow' into integrate 2026-10-03 23:12:03 +02:00
jochen 980a0dee93 integrate: 214 2026-10-03 23:12:03 +02:00
jochen 6784efae75 A commit is never a branch to follow (hq issue 215)
A build asked at a commit recorded that commit as the module's ref. Every merge after it failed to
match the module and its plan left it out without a word, and every plan that rebuilt it asked for
the same old commit again. Registration now keeps the branch the module followed (the default
branch for a new one); matching and re-asking read a recorded commit as the default branch, which
heals records already pinned this way; and a merge says which modules of its repository it leaves
out because they follow another branch.
2026-10-03 22:22:08 +02:00
jochen d86baebe9a A plan settles an asked build from the build records (hq issue 214)
A merge to the controller's own repository replaces the controller in its first tier; the build
that produced the new one was recorded, the plan never heard it, and it waited for ever with every
later plan behind it. The record is the fact: a build recorded after the ask is the tier's outcome,
whoever was listening when it came.
2026-10-03 22:20:33 +02:00
jochen 58b4fcb8c8 A bundle stands on the toolchain it is compiled in (hq issue 211)
A manifest names its toolchain by language, not in build.on, so the planner did not know a bundle
depends on the module that publishes its toolchain and built the two in one tier: the bundle
against the old toolchain, recorded as built from the new commit. The edge is read from the
manifest, so it holds before any build recorded it, and a toolchain that moves rebuilds every
bundle compiled in it.
2026-10-03 22:18:20 +02:00
jochen e67c58cd98 Every serving principal may answer the services discovery for what it serves; the controller announces its seat (hq ADR 0197)
Grants: a principal that serves tools subscribes $SRV.PING/$SRV.INFO and those questions under
each name it serves — its own and no other's; the tool runtime and people may ask. The controller
answers discovery for the mesh-controller seat in NATS's services format, one endpoint per verb it
serves, with the seat's description and schema. module list --json says which modules declare tools,
so the console expects an announcement only from those.
2026-10-03 22:11:00 +02:00
jochen 85873b19e1 node list and module list answer --json, and the nodes and modules verbs use it (hq ADR 0195)
The console's discovery reads the machines and the modules; parsing a printed column breaks when it
is reworded. Both now answer JSON on --json, as status and seats do, and the seat verbs ask for it.
2026-10-03 21:55:35 +02:00
jschoubben 11e4bc0ba1 The roster is the machines: each node's internal domain covers its routes (hq ADR 0191)
The roster published routed names — public ones first, then (in this PR's first take) internal ones
told apart by suffix. Neither is needed: a node has one internal domain and every route on it is a
name under it, answered by the resolver's per-node wildcard; a node's public domains are public
DNS's. routeNamesInTheMesh and NamesServed are removed, and a test pins .Names to the machines.
2026-10-03 16:10:16 +02:00
jschoubben e56f3aa1cb The roster publishes a route's internal name, never its public one (hq ADR 0191)
NamesServed read a route's public `name` and plan.go then filtered by suffix — telling the mesh's
names from public ones by their spelling, when the mesh composed both itself. It now publishes the
`internal-name` it composed under the serving node (ADR 0151); the suffix filter is gone.
2026-10-03 15:39:05 +02:00
jschoubben 408f6dbad9 The mesh answers only its own names privately; a public name resolves publicly (hq ADR 0191)
Every routed public name was published into each machine's hosts region at its serving node's
private address. ace's resolver also answers its LAN, so a phone there got the control-node's
tunnel address for the mail server and could not connect. Routes have internal names under the
serving node (ADR 0151), so only names under the mesh suffix are published now.
2026-10-03 15:14:01 +02:00
mesh-admin eae0577567 Merge pull request 'Issues 203 and 206: an assignment issues its credential; the controller owns a worker's shape; the build seat's holder follows the controller' (#233) from fix/issues-203-206 into main 2026-10-03 09:49:54 +00:00
jochen 28853a251b The build seat's holder follows the controller that defines its worker (hq issue 206)
A plan is ordered by artifacts and says nothing about what must be running before what (ADR 0162);
on 2026-10-03 that put the build machine in tier 0 and the controller in tier 1, and the new build
machine could not bind the worker the old controller had defined. One running order enters the
graph, named as its own edge: a module claiming the build seat follows the control plane, and the
built-by edge from the control plane to that holder yields to it — the controller is built by
whichever build machine is running, as the runtime image always was. The edge orders a plan and
never widens it, like built-by.
2026-10-03 04:04:41 +02:00
jochen f86f6a74f0 A declaration is numbered when it is composed, and a send is recorded even by a sender being replaced (hq issue 204)
On 2026-10-02 a runtime assigned and applied on two machines was undone two seconds later by a
declaration that had the assignments of a minute earlier. Every path composes from the records at
compose time and holds the machines it sends — but the number went on at SEND time, after
composing, so a declaration composed before an assignment changed and sent after a newer one
carried the higher number, and the host, which rightly refuses a lower number, took the older
content as the mesh's newest word. The record of that send was never written either: it is written
after the declaration is away, on the sender's context, and the controller sending it was being
replaced in that very second — status read "applied, current" over a machine just told otherwise.

Now the number is taken before the composition reads anything, in every path, so what was composed
earlier is numbered lower however late it goes out and the host's refusal does what it is for; and
what was sent is written down on a context that outlives the sender, bounded, so a dying controller
still records what it told a machine. The `declare` command — a declaration a person sends by hand —
records its send too. Proven: compositions in one order and sends in the other keep the numbers in
composition order; a send is recorded after the sender's context is cancelled.
2026-10-03 04:02:36 +02:00
jochen a3e8a4185b An assignment issues its bus credential, a push refuses one nobody issued, and what reads a secret restarts on it (hq issue 203)
`assign` recorded a module and `push` sealed a random own secret where its bus credential belongs;
the process crash-looped until a person ran `module issue` and pushed again, and the only warning was
one line in a list printed on every push. Now assigning a module that declares a broker secret issues
the credential in the same act — kept when one exists, so re-assigning rotates nothing — and when the
bus cannot be reached from here the assignment says which verb to run. A push never seals a
placeholder in a credential's place: a module whose bus user is unminted is refused by name, with the
verb. The control plane's own user is the installer's, seeded at genesis, which the test now says.

And what reads one of a module's own secrets is restarted when it changes — composed for a container
or daemon that names the secret's path in its volumes, environment or env-files, so a manifest need
not say it: the build machine ran on an hour-old credential because its manifest restarted it on its
environment file alone (issue 206). A scheduled or run-once process is left alone; it reads afresh.
2026-10-03 04:01:17 +02:00
mesh-admin 78de54381b Merge pull request 'A seat's work is shared by its holders: node-build-agent, pulled one ask at a time (hq ADR 0190)' (#228) from feat/a-seats-work-is-shared-by-its-holders into main 2026-10-03 00:53:42 +00:00