Two different things were both written `requires`. A shell, a display
server and a private network have to be on the machine that needs them.
A database does not — it runs somewhere and is reached over the network.
Both were answered the same way, so requiring a database installed
PostgreSQL on every machine that ran a web application.
What a module provides now carries a scope, the same idea claims already
use, written short in the ordinary case:
"provides": ["shell"]
"provides": [{"name": "database", "scope": "mesh"}]
A mesh-scoped requirement is answered by finding the node already running
it — never by installing it here. Choosing a machine to put a database on
is a decision with consequences, and nothing resolving a web application
should make it silently. With nothing anywhere it refuses and says which
module to assign; with two it refuses and says how to choose.
Choosing is `pin <node> <provision> <from>`, kept per node because that
is the granularity the choice has. A pin at a machine that does not
provide it refuses rather than falling back — a fallback would quietly
move somebody's data. One provider does not overrule a pin either.
Resolving a node now needs to know what the others offer, and working
that out needs them resolved, so it is two passes: the first answers only
what each node offers, the second answers everything. Nothing is ever
declared from the first.
A node's plan says what it takes from elsewhere. It is the only part of a
set that stops working when a different machine goes away, and nothing
else in that output would have said so. It is also where a credential
will hang once there is a mechanism for handing one back.
One test found passing for the wrong reason: it read pins through a join
on the provider, which hides a dangling row whether or not it was cleaned
up. It counts rows now, and bites when the cascade is removed.
Delivery is a comparison, not a pipeline: the control plane holds what source
exists and what has been built from it, and the difference is the work. Both
halves are written down now, so "is this current" is a question about two
columns rather than something you find out by building.
`status` answers "did my change go out?", which ADR 0010 names as the real risk
of replacing a pipeline with a comparison -- it is answerable today by opening
a pipeline, and something had to replace that.
zsh holds 4f2a9c1e, source has 9e3b7d2a
running on laptop
The machines are the point. A module being out of date is a fact about the
catalogue; which machines are running last week's version is the thing with
consequences.
Three things this had to get right.
A module with no source is never behind -- it was handed over directly, which
is how a one-off arrives, and saying "out of date" about it would be inventing
a comparison against nothing.
A source nobody has checked is not behind either. Reporting it as behind would
put every module on the list the moment provenance was recorded, which makes
the list say nothing. Fault injection found this: my first test passed with the
guard removed, because both halves were empty strings and compared equal. The
case that actually needed it -- a known commit and an unknown head -- was
untested.
And handing over a manifest by hand does not erase where the module normally
comes from. Fixing something in a hurry is legitimate; silently forgetting its
origin is not, because that record is the only thing that would say afterwards
that a machine is running something nobody can rebuild.
Also fixed the flag parsing, which stopped at the first positional argument and
silently ignored every flag after it -- so `module add thing.json --source x`
recorded no source at all and said it had succeeded. The host's own parser
documents this exact footgun and I wrote it again anyway.
The half of the module system that was built and never proved. Unassigning i3
removed i3's file AND xorg's, because xorg was only there to satisfy i3 -- the
node's own record agrees, and the resolution the mesh sends no longer mentions
either.
That works because a declaration removes what the mesh previously declared and
nothing else, which is 04-ISSUES/010's fix carrying its weight here: the
substrate the machine raised for itself is untouched by any of it.
Tests for the storage layer, which had none. The ones worth naming:
A module a machine is running cannot be forgotten -- not a fault, it means the
mesh would lose the ability to describe what is on that machine. Removing a
node DOES take its assignments, and the asymmetry is deliberate: a node that is
gone cannot be running anything.
A node that has never reported has NO capabilities rather than all of them.
That refuses anything needing one, which is wrong but visible -- where assuming
it can do everything would assign work it cannot do and find out on the
machine. And a capability the node reported as ABSENT is not counted: reading
the list without the verdict would let a module onto a machine that said no.
`overlay push` is gone, replaced by `push`, which sends a node its network and
its modules as one declaration. Two commands that overlap is how a mesh ends up
half-configured by whichever was run. The old name answers with where to go,
and answers before opening a database -- needing one would turn a redirect into
a connection error.