status reported no unmet dependency on any node once systemd, pacman and docker
were assigned to all four (to-be 42), which is the condition ADR 0207 set for the
switch. A dependency no catalogue module could meet stays a report before and
after the switch, as assign already said it: there is no remedy to name.
assign and unassign say only what they changed on their node; push <node>
lists that node's, push to many counts each and points at status. The
once-per-change log is the serving controller's alone: a one-shot command
starts with no memory, so it logged every node on every call.
The host only ever adds groups, so the container runtime's module can put the
operator in its group while the shell's module sets the same account's shell.
Seed the eleven node seats with the verbs they start with. A provision may
have the machine's reach: a requirement for it resolves only to a provider
in the node's own set, is never pulled in, and is refused naming who could.
A shell contribution's for gains xinitrc and xresources, placed only by the
holder of node-display-server.
Seed node-package-manager and node-container-runtime. Derive each module's
dependencies from its declared service, package and container resources;
judge them over the node's whole set, exempting the foundation. Refuse at
assign (several modules may go on as one act) and at unassign of the last
holder; report at composition in status, behind one switch.
The short form is a question the mesh answers: "80" means publish what the
software calls 80, and the mesh fills in the machine's half from the port it
assigned. It can only assign one for a port the module declared, so a number
appearing nowhere in listens gets no assignment and reaches the machine as
written — which is how the photo module asked for port 80 on the node whose
reverse proxy holds it.
Four modules publish 80 quite safely, because they declare 80. The difference
is the declaration, not the number. A catalogue-wide test now says so; it
names all three offenders against the catalogue as it was.
Issue 225. The mesh seals one credential per consumer beside the provider's
contributions file, and wrote it root-owned. That was right while a module's
own code ran in a container as root; ADR 0198 moved that code under the node's
runtime, as the node's account, and the secret stayed root's. On the control
machine two consumers went unprovisioned for three hours and the only sign
was a line reading 'secret not readable yet', 4330 times.
The same sentence is already written for a module's own secrets a few hundred
lines above — 'a root-owned 0600 file is one that process cannot read'. This
is that rule reaching the other kind of secret the mesh writes for a module.
Issue 226. The sweep met a reference recorded with the store's old address,
read 'I will not address this' as 'the store refuses everything', and
collected none of the 1681 it had found. Two changes: references from build
records are read through Recorded, where the provenance is known — not in
LetGo, which cannot tell one registry host from another and must stay strict
— and a reference the sweep will not address is now ErrNotOurs, skipped,
never a reason to stop. Only the store refusing ends a sweep.
make check: the two failures both fail on main as well — the resolver test
(hq 202/203) and the service-manager test, which reads this machine's own
shell environment.
The raise at start was the only place buckets were asserted, so a module
registered and assigned since had none until the control plane restarted —
found on the first module to declare state.
A module names its own resources locally; a declaration names them under the
module. restart-on and reload-on are rewritten for exactly that reason and
while-stopped was not, so the store's step said it held "store" still while
the machine's container is "distribution.store".
The host refuses a declaration naming a container it does not have — whole.
So novox took nothing at all, on every push, from 04:15 until this. The
machine was never damaged: refusing whole is what kept it serving.
Both sides' tests passed throughout. The controller's read manifests, the
host's read hand-written declarations with bare ids, and nothing composed one
and judged the result. That test now exists.
A module contributes environment variables, PATH entries and shell code in named slots;
the holder of the matching seat places them with ${environment:posix|systemd} and
${shell:<shell>:<slot>}. Rendered in module order with a naming line per contribution,
PATH entries added only when missing, machine facts resolved first. A variable two
modules set, or a placeholder outside its seat's holder, is refused at parse (the
catalogue check) and at composition. Filled after every other placeholder pass, so no
scanner ever reads a shell's own ${...}.
node-environment says which module writes the account's environment; node-login-shell
replaces the module-declared login-shell, so a second shell claims it rather than
declaring a rival, and execute is the mesh's contract. login-shell is refused as a
module's seat name. Seeded into a live store by the existing additive seeding.
Three things found reading this back, each of which would have been quiet.
A consumer that keeps several holders of one provision (ADR 0094) gets a
login per holder, and a provider derives from the login — so it would make a
resource per holder while the consumer is told one value for the requirement.
That is issue 124's own failure one case to the side: authenticate, then be
refused on every object. Refused now, naming both ends.
The sweep runs inside somebody's build and was unbounded. At most two hundred
artifacts and sixty seconds, stopping at the first refusal because a store
that refuses one refuses all; the rest is offered again next build.
The citation and migration renumbers are in the commit before this one.
The bundles refactor took ADR 0188 on main, so this work's record is 0201 and
every comment citing it moves with it. Main also took migration 0055 (an
older build never replaces a newer), so the store's collected-artifacts table
is 0056 — a number two migrations share is a schema nobody can trust.
make check passes except TestTheResolverIsToldEveryMachineOnTheNetworkAndToldAgainWhenOneLeaves,
which fails on main too and now for two stacked reasons (hq issues 203 and 202).
A manifest names the state it keeps (state) and reads (reads); the controller
asserts a key-value bucket per name on every raise, grants owners write and
readers read (measured against a running server), issues each assignment its
buckets in the membership, and reports buckets nothing declares without
removing them.
The mesh names what may go from its own build records — a digest it did not
record making is never named, which is what keeps the sweep away from the
images genesis pushed. An artifact stays because a definition the mesh holds
names it, or because it belongs to one of the five most recent successful
builds of its module.
internal/artifacts asks the store to let go of one; internal/inventory
decides and remembers (migration 0055); the sweep runs after a build the mesh
recorded, which is when both the bytes and the keep set moved. Never fatal to
a build.
And the manifest side of while-stopped, refused from the definition alone:
no schedule, run-once, a container the module does not declare, itself.
${consumer:as} and ${consumer:as:dns} in a serves block are filled per
consumer at resolution, and the one filled value reaches both ends: the
consumer's binding and its ${bound:...} substitutions, and the provider's
contributions entry as `derived`. A fact or alphabet the mesh does not have
is refused at parse; a consumer whose own file already holds the derived
value is refused at resolution, naming the placeholder to write instead.
While some thirty modules still stood in that shape, one already registered so was rebuilt without
complaint. Every module has moved since; the exception would only let one move back.
Genesis now raises a process-form controller as a container built from
this repository's Dockerfile with no build arguments (mesh-host
bootstrap, novox/hq issue 223); the manifest builds no image, so nothing
passes the base in. The default was a tag older than go.mod asks for.
It is now the digest the Makefile pins, and a test holds the two equal.
The controller is a Go program and was the one piece of the mesh's own Go
code still shipped and run as an image (novox/hq issue 213; ADR 0188 §1:
a module's own code is bundles; §3: a service bundle is a process).
The manifest now builds one Go bundle, `controller`, and runs it as the
process `mesh-controller` (`./mesh-controller serve`) under an account
the module declares. What the container gave it, replaced:
- host network: a process is on the host's network; nothing it reads
names a container network
- user 65534: the account `mesh-controller`, which owns its secrets and
its state directory
- the eight mounts: the env names the host paths the mesh already places
(the store, broker and bus files under the state directory, the
broker's certificate under /var/lib/mesh-broker-tls); the `broker`
mount was read by nothing and is gone with the others
- `container-runtime` is no longer required on its machine
Its preparation is the same binary with `prepare`, as a run-once process,
and the process `replaces` the container `server`: the host keeps the
container answering until the process is running (mesh-host). Needs the
previous commit live in the running controller, and the host's
`replaces` on the controller's machine, before it is registered.
No image is built by the mesh any more. The Dockerfile stays for genesis
and the lab (`make image`, its Go base now pinned in the Makefile).
A bundle could import only what the toolchain image carried: the compiler and the bundler resolve an import from the module's directory and then the toolchain's node_modules, and nothing ever put anything in the first. So a module needing a database driver (pg, mongodb, mssql) could not be a bundle, and kept a container whose recipe installed it (hq ADR 0198 §4: the backend's own driver inside the bundle).
Now, when a module's package.json depends on anything beyond the SDK, the build installs its production dependencies into the module's directory, in the toolchain image, before the compile: npm ci from the lockfile when there is one, npm install from the ranges otherwise, the mesh's registry for the SDK's scope and the public one for the rest, install scripts off. esbuild then inlines them. A module depending only on the SDK runs exactly the commands it did before.
The SDK stays the toolchain's (hq issue 212): it is taken out of what is installed and any copy something pulls in is removed, so every import of it resolves past the module's node_modules to the one the toolchain carries; a module's own range never shadows it. npm's verified download cache is a named volume; nothing installed is kept between builds. Without a registry, a scoped package is refused rather than resolved on the public registry.
The controller's machine moves it from the container to a process by
starting the process first and removing the container once the process
is up (mesh-host's `replaces`). For that moment two controllers share the
store and the bus. Checked what each does:
- the seat's verbs: a queue group per seat, each call answered once. Safe.
- the controller's consumers on CONTROL and EVENTS: push consumers with
no delivery group, so the second bind is refused with "consumer is
already bound" and serve exited. The process would restart for ever,
the host would never see it up, and the container would never go. The
second controller now stands by and binds when the first lets go
(tested on a real bus; fails without the change).
- plans: read, changed and saved whole by the 30s timer, by build
outcomes, by a merge and by `plans stop`. Two timers would each ask a
tier the other had just asked. Working the plans now takes a
session-level advisory lock on the inventory: the timer skips while
another holds it, the other paths wait for it. Build asks happen only
inside plan work and are covered by the same lock.
The controller is to be declared as a Go bundle run by a process instead of
an image (novox/hq issue 213, ADR 0188 §1, §3). The composer could not
express that honestly yet:
- a module declaring tools had every bundle served by the node's runtime,
so the controller's own binary would have been launched a second time as
an MCP child; a bundle one of the module's resources runs is now served
only when it says `loads`
- a module's accounts went after the mesh-computed files, so secrets owned
by the account a process runs as were refused on the first apply; a
module's `user` resources now go first
- `prepares` derived its step only from a container; a process is now
prepared by the same program with `prepare` as a run-once process
- a process may say what it `replaces` (a resource of its module it no
longer declares), prefixed as the host records it, so the host keeps the
old one running until the process is (needs mesh-host's `replaces`)
This lands before the controller's manifest uses any of it: the running
controller composes its own declaration, so the code that fills the new
shape must be live first.
gitea's own code moves out of its runtime container (mesh-catalog, to-be 38 WP4c waves 2-3), so the three tests that composed the forge from the catalogue beside this checkout resolve its build as the code bundle, compose it beside the node's runtime, and read the forge's address from the words the runtime hands the module rather than from a sidecar's env.
A module's own code moving out of its container (novox/hq to-be 38 WP4c)
becomes a process on the machine, and still has to be told what its
container was: the port this machine gave the module and where the
foundation's seats are. ${port:…} and ${seat:…} were filled only in a
file's content and a container's env, so in a process's env they reached
the machine as literals, and the modules that moved first (mesh-catalog
#245) wrote their run-once steps a 0600 env file instead. A process's env
now takes the same resolution and the same refusals; ${dir:…} and
${access:…} already did, and a bundle's env (ADR 0192) already resolves
${dir:…} and ${port:…}.
Builds of one module in flight together finish in any order, and the mesh
took whatever it heard last as what the module is: RegisterModule overwrote
the module's manifest unconditionally, and Held/BuiltAgainst/ReadRepositories
ordered builds by when they were recorded. A postgres build asked before the
mesh-tools runtime fix finished after the one asked after it, and the next
push deployed the stale image (novox/hq issue 219).
A build is now ordered by when it was asked, read from the build-<nanos> id
the controller writes: build.asked and module.built_asked (migration 0055).
A registration from an earlier request than the module's current one is
recorded and refused as superseded. A plan takes as its outcome only a build
asked at or after its own ask, so an earlier plan's leftover build cannot
settle a later plan. Ids of any other shape keep the old order.