cd2481dcd89e9ffe002d1093a2bb355932d4e6e7
8
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
97448194ac |
Seats are a closed set, a seat's holder answers for what it delivers, and a build source may live on the git seat
Implements novox/hq ADR 0110 and 0111. The seat set lives in internal/catalogue/seats.go: fourteen seats, each with a scope, what occupying it delivers, and the record that made it one. A test asserts the count and a decision per entry, so changing the set means finding the argument, as the host's vocabulary test does. The first set is every seat already claimed — including the-private-network, which the network module claims from a manifest composed in this repository's code, not from any module.json — plus npm-package-registry (ADR 0109) and git (ADR 0111). A test parses every catalogue manifest and this repository's own and fails on any refused claim, so closing the set refuses nothing in use. ParseManifest now refuses a claim on a seat the mesh does not define, a seat claimed at another scope, and a delivering seat claimed by a module that does not provide what it delivers. A malformed claim is refused once, for being malformed. Resolution: among several providers of a mesh provision, a pin still wins; then the holder of the seat that delivers it; then the only provider; otherwise refused as before. ADR 0009's "never guessed" holds — the seat is the choice made once, mesh-wide, rather than a pin per consumer node. A provider now carries the module it came from, because a provider is a (node, module) pair and the pair is what tells a holder from a neighbour on the same machine. The planner's second pass is now given the first pass's holdings. Without them, a node consuming a seat-delivered provision was refused there, and a refused node's own claims dropped out of what the mesh holds — letting a second holder of one of its seats pass unrefused. `seats [--json]` lists every seat, what it delivers, and each holder, derived from assignments every time and never stored. Unheld seats are listed. A stored claim outside the set — possible for a manifest registered before the set closed, since stored manifests are not re-validated — is shown rather than hidden. `build --self <owner>/<repo>` builds from a repository on the git seat's holder. The clone URL is composed at build time from the holder's node and what it serves for git; the recorded source is the path and the seat (migration 0032), never an address, so a moved forge changes nothing recorded. Nobody holding the seat refuses self-hosted builds and says so; external URLs are unchanged. An address passed with --self is refused rather than recorded as a path. Replaces three foundation tests that defended the builder's carried package binding. The catalogue removed that binding when the builder began requiring the registry through a real grant, so the tests were already failing on main; they now assert the builder requires what the npm seat delivers and carries no copy of its own, and that the forge holds the npm and git seats. Verified: go vet clean; the whole suite passes against a throwaway Postgres (make postgres), the new inventory tests included; gofmt clean apart from cmd/mesh-builder/stdout_test.go, which fails on main too. |
||
|
|
c3b88b9148 |
Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo becomes mesh-controller, the seat the-controller, and the store+broker pair the foundation (embedded base bundles, default template and example lock renamed with their go:embed directives). No behaviour change — a pure vocabulary rename. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx |
||
|
|
ee84b624b1 |
A provision names the engine, because a consumer is coupled to one
Provisions were named after roles: provides "database", requires "database". Nothing distinguished engines, so a module written against PostgreSQL could be matched to a provider of SQL Server, resolve as satisfied, deploy, and fail on its first query — with nothing connecting that error back to a match made elsewhere by something that believed it had done its job. The failure is in the direction that hides. Refusing on ambiguity exists precisely so this does not happen, and the generic name walked around it: with one provider of each name nothing is ambiguous, so nothing is asked. How it got in: every resolver test had exactly one provider per name, so no mismatch was expressible and none was caught. The fixtures agreed with the design — the same fault as the imagined test output in 04-ISSUES/005, at the level of a name. Refused rather than documented, because the old naming *was* the documented convention. Providing database/db/sql/sql-database is now a parse error naming what to write instead. The rule is about coupling, not specificity everywhere: route and resolver stay role-named, because a consumer genuinely cannot tell which proxy answered. novox/hq ADR 0027. |
||
|
|
dfca21fa55 |
Keep what a machine said about itself, not just the yes
novox/hq ADR 0009: a capability's presence gates an assignment and its detail carries a value — seat: card1-DP-1, an architecture, an amount of memory. So 'can this run here' and 'what should it be configured as' are one fact read two ways, and the mesh was keeping the first read and discarding the second. The reason an absent capability is absent went the same way, which is the case a person most needs: 'this machine has no container runtime' is the answer and 'docker is not installed' is why, and only the machine knows why. `node show` says it back. Never reported and reported nothing stay different things there — one machine has not run the host, the other ran it and can do nothing, and those send a person to different places. |
||
|
|
421fe73dce |
The mesh builds: a machine takes the work, and the catalogue shows it
A build is work, not state. Everything else the control plane sends a node is a declaration — this is what you should be — reconciled forever. A build happens once and is finished. Putting it in a declaration would mean rebuilding on every reconcile, or a declaration carrying "and I already did this", which is state about an event rather than about a machine. So it travels on its own queue and the answer comes back correlated. One queue, so several build machines share the work and each request is done exactly once — which a per-machine routing key would not give. mesh-builder is the program a build machine runs. Not the control plane, which must not run commands on a machine; not the host, which would then need a container runtime and git everywhere to do something almost no machine will ever do. It holds its own broker credential and nothing else. Three properties that are decisions: - a request is acknowledged only once the answer is away, so a builder that dies mid-build leaves the work for another machine rather than losing it with nobody ever hearing why - one build at a time. Five at once against one runtime finishes all five slower than it would have finished the first, and the queue is what shares work between machines - a failure is a RESULT. A build that fails silently is indistinguishable from a builder that is not running, and those want different responses And `module list` is a catalogue: what exists, at which version, built from which commit or handed over by hand or shipped with the control plane, whether it is behind its source, and which machines run it. All of that was recorded from the first build and none of it was shown, so "is this current?" could only be answered by reading the database. Proven against a real broker, registry and store: the mesh asked, a builder consumed, built, published, answered; the manifest was recorded with its commit; the source moved and the catalogue said "behind"; rebuilding caught it up with a new digest because the content changed. |
||
|
|
d4064122d6 |
Where the answer to a requirement is allowed to live
Two different things were both written `requires`. A shell, a display
server and a private network have to be on the machine that needs them.
A database does not — it runs somewhere and is reached over the network.
Both were answered the same way, so requiring a database installed
PostgreSQL on every machine that ran a web application.
What a module provides now carries a scope, the same idea claims already
use, written short in the ordinary case:
"provides": ["shell"]
"provides": [{"name": "database", "scope": "mesh"}]
A mesh-scoped requirement is answered by finding the node already running
it — never by installing it here. Choosing a machine to put a database on
is a decision with consequences, and nothing resolving a web application
should make it silently. With nothing anywhere it refuses and says which
module to assign; with two it refuses and says how to choose.
Choosing is `pin <node> <provision> <from>`, kept per node because that
is the granularity the choice has. A pin at a machine that does not
provide it refuses rather than falling back — a fallback would quietly
move somebody's data. One provider does not overrule a pin either.
Resolving a node now needs to know what the others offer, and working
that out needs them resolved, so it is two passes: the first answers only
what each node offers, the second answers everything. Nothing is ever
declared from the first.
A node's plan says what it takes from elsewhere. It is the only part of a
set that stops working when a different machine goes away, and nothing
else in that output would have said so. It is also where a credential
will hang once there is a mechanism for handing one back.
One test found passing for the wrong reason: it read pins through a join
on the provider, which hides a dangling row whether or not it was cleaned
up. It counts rows now, and bites when the cascade is removed.
|
||
|
|
653e232f1c |
The mesh knows where a module came from, and whether it is behind
Delivery is a comparison, not a pipeline: the control plane holds what source
exists and what has been built from it, and the difference is the work. Both
halves are written down now, so "is this current" is a question about two
columns rather than something you find out by building.
`status` answers "did my change go out?", which ADR 0010 names as the real risk
of replacing a pipeline with a comparison -- it is answerable today by opening
a pipeline, and something had to replace that.
zsh holds 4f2a9c1e, source has 9e3b7d2a
running on laptop
The machines are the point. A module being out of date is a fact about the
catalogue; which machines are running last week's version is the thing with
consequences.
Three things this had to get right.
A module with no source is never behind -- it was handed over directly, which
is how a one-off arrives, and saying "out of date" about it would be inventing
a comparison against nothing.
A source nobody has checked is not behind either. Reporting it as behind would
put every module on the list the moment provenance was recorded, which makes
the list say nothing. Fault injection found this: my first test passed with the
guard removed, because both halves were empty strings and compared equal. The
case that actually needed it -- a known commit and an unknown head -- was
untested.
And handing over a manifest by hand does not erase where the module normally
comes from. Fixing something in a hurry is legitimate; silently forgetting its
origin is not, because that record is the only thing that would say afterwards
that a machine is running something nobody can rebuild.
Also fixed the flag parsing, which stopped at the first positional argument and
silently ignored every flag after it -- so `module add thing.json --source x`
recorded no source at all and said it had succeeded. The host's own parser
documents this exact footgun and I wrote it again anyway.
|
||
|
|
931a3a19a5 |
Taking a module off a node takes it off the machine
The half of the module system that was built and never proved. Unassigning i3 removed i3's file AND xorg's, because xorg was only there to satisfy i3 -- the node's own record agrees, and the resolution the mesh sends no longer mentions either. That works because a declaration removes what the mesh previously declared and nothing else, which is 04-ISSUES/010's fix carrying its weight here: the substrate the machine raised for itself is untouched by any of it. Tests for the storage layer, which had none. The ones worth naming: A module a machine is running cannot be forgotten -- not a fault, it means the mesh would lose the ability to describe what is on that machine. Removing a node DOES take its assignments, and the asymmetry is deliberate: a node that is gone cannot be running anything. A node that has never reported has NO capabilities rather than all of them. That refuses anything needing one, which is wrong but visible -- where assuming it can do everything would assign work it cannot do and find out on the machine. And a capability the node reported as ABSENT is not counted: reading the list without the verdict would let a module onto a machine that said no. `overlay push` is gone, replaced by `push`, which sends a node its network and its modules as one declaration. Two commands that overlap is how a mesh ends up half-configured by whichever was run. The old name answers with where to go, and answers before opening a database -- needing one would turn a redirect into a connection error. |