Commit Graph
7 Commits
Author SHA1 Message Date
jschoubben c3b88b9148 Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben ee84b624b1 A provision names the engine, because a consumer is coupled to one
Provisions were named after roles: provides "database", requires
"database". Nothing distinguished engines, so a module written against
PostgreSQL could be matched to a provider of SQL Server, resolve as
satisfied, deploy, and fail on its first query — with nothing
connecting that error back to a match made elsewhere by something that
believed it had done its job.

The failure is in the direction that hides. Refusing on ambiguity
exists precisely so this does not happen, and the generic name walked
around it: with one provider of each name nothing is ambiguous, so
nothing is asked.

How it got in: every resolver test had exactly one provider per name,
so no mismatch was expressible and none was caught. The fixtures agreed
with the design — the same fault as the imagined test output in
04-ISSUES/005, at the level of a name.

Refused rather than documented, because the old naming *was* the
documented convention. Providing database/db/sql/sql-database is now a
parse error naming what to write instead.

The rule is about coupling, not specificity everywhere: route and
resolver stay role-named, because a consumer genuinely cannot tell
which proxy answered. novox/hq ADR 0027.
2026-08-31 17:12:46 +02:00
jschoubben dfca21fa55 Keep what a machine said about itself, not just the yes
novox/hq ADR 0009: a capability's presence gates an assignment and its detail
carries a value — seat: card1-DP-1, an architecture, an amount of memory. So
'can this run here' and 'what should it be configured as' are one fact read two
ways, and the mesh was keeping the first read and discarding the second.

The reason an absent capability is absent went the same way, which is the case
a person most needs: 'this machine has no container runtime' is the answer and
'docker is not installed' is why, and only the machine knows why.

`node show` says it back. Never reported and reported nothing stay different
things there — one machine has not run the host, the other ran it and can do
nothing, and those send a person to different places.
2026-08-31 04:53:45 +02:00
jschoubben 421fe73dce The mesh builds: a machine takes the work, and the catalogue shows it
A build is work, not state. Everything else the control plane sends a
node is a declaration — this is what you should be — reconciled forever.
A build happens once and is finished. Putting it in a declaration would
mean rebuilding on every reconcile, or a declaration carrying "and I
already did this", which is state about an event rather than about a
machine.

So it travels on its own queue and the answer comes back correlated. One
queue, so several build machines share the work and each request is done
exactly once — which a per-machine routing key would not give.

mesh-builder is the program a build machine runs. Not the control plane,
which must not run commands on a machine; not the host, which would then
need a container runtime and git everywhere to do something almost no
machine will ever do. It holds its own broker credential and nothing
else.

Three properties that are decisions:

- a request is acknowledged only once the answer is away, so a builder
  that dies mid-build leaves the work for another machine rather than
  losing it with nobody ever hearing why
- one build at a time. Five at once against one runtime finishes all five
  slower than it would have finished the first, and the queue is what
  shares work between machines
- a failure is a RESULT. A build that fails silently is
  indistinguishable from a builder that is not running, and those want
  different responses

And `module list` is a catalogue: what exists, at which version, built
from which commit or handed over by hand or shipped with the control
plane, whether it is behind its source, and which machines run it. All of
that was recorded from the first build and none of it was shown, so "is
this current?" could only be answered by reading the database.

Proven against a real broker, registry and store: the mesh asked, a
builder consumed, built, published, answered; the manifest was recorded
with its commit; the source moved and the catalogue said "behind";
rebuilding caught it up with a new digest because the content changed.
2026-08-30 03:46:02 +02:00
jschoubben d4064122d6 Where the answer to a requirement is allowed to live
Two different things were both written `requires`. A shell, a display
server and a private network have to be on the machine that needs them.
A database does not — it runs somewhere and is reached over the network.
Both were answered the same way, so requiring a database installed
PostgreSQL on every machine that ran a web application.

What a module provides now carries a scope, the same idea claims already
use, written short in the ordinary case:

  "provides": ["shell"]
  "provides": [{"name": "database", "scope": "mesh"}]

A mesh-scoped requirement is answered by finding the node already running
it — never by installing it here. Choosing a machine to put a database on
is a decision with consequences, and nothing resolving a web application
should make it silently. With nothing anywhere it refuses and says which
module to assign; with two it refuses and says how to choose.

Choosing is `pin <node> <provision> <from>`, kept per node because that
is the granularity the choice has. A pin at a machine that does not
provide it refuses rather than falling back — a fallback would quietly
move somebody's data. One provider does not overrule a pin either.

Resolving a node now needs to know what the others offer, and working
that out needs them resolved, so it is two passes: the first answers only
what each node offers, the second answers everything. Nothing is ever
declared from the first.

A node's plan says what it takes from elsewhere. It is the only part of a
set that stops working when a different machine goes away, and nothing
else in that output would have said so. It is also where a credential
will hang once there is a mechanism for handing one back.

One test found passing for the wrong reason: it read pins through a join
on the provider, which hides a dangling row whether or not it was cleaned
up. It counts rows now, and bites when the cascade is removed.
2026-08-29 23:51:50 +02:00
jschoubben 653e232f1c The mesh knows where a module came from, and whether it is behind
Delivery is a comparison, not a pipeline: the control plane holds what source
exists and what has been built from it, and the difference is the work. Both
halves are written down now, so "is this current" is a question about two
columns rather than something you find out by building.

`status` answers "did my change go out?", which ADR 0010 names as the real risk
of replacing a pipeline with a comparison -- it is answerable today by opening
a pipeline, and something had to replace that.

  zsh    holds 4f2a9c1e, source has 9e3b7d2a
         running on laptop

The machines are the point. A module being out of date is a fact about the
catalogue; which machines are running last week's version is the thing with
consequences.

Three things this had to get right.

A module with no source is never behind -- it was handed over directly, which
is how a one-off arrives, and saying "out of date" about it would be inventing
a comparison against nothing.

A source nobody has checked is not behind either. Reporting it as behind would
put every module on the list the moment provenance was recorded, which makes
the list say nothing. Fault injection found this: my first test passed with the
guard removed, because both halves were empty strings and compared equal. The
case that actually needed it -- a known commit and an unknown head -- was
untested.

And handing over a manifest by hand does not erase where the module normally
comes from. Fixing something in a hurry is legitimate; silently forgetting its
origin is not, because that record is the only thing that would say afterwards
that a machine is running something nobody can rebuild.

Also fixed the flag parsing, which stopped at the first positional argument and
silently ignored every flag after it -- so `module add thing.json --source x`
recorded no source at all and said it had succeeded. The host's own parser
documents this exact footgun and I wrote it again anyway.
2026-08-29 22:32:16 +02:00
jschoubben 931a3a19a5 Taking a module off a node takes it off the machine
The half of the module system that was built and never proved. Unassigning i3
removed i3's file AND xorg's, because xorg was only there to satisfy i3 -- the
node's own record agrees, and the resolution the mesh sends no longer mentions
either.

That works because a declaration removes what the mesh previously declared and
nothing else, which is 04-ISSUES/010's fix carrying its weight here: the
substrate the machine raised for itself is untouched by any of it.

Tests for the storage layer, which had none. The ones worth naming:

A module a machine is running cannot be forgotten -- not a fault, it means the
mesh would lose the ability to describe what is on that machine. Removing a
node DOES take its assignments, and the asymmetry is deliberate: a node that is
gone cannot be running anything.

A node that has never reported has NO capabilities rather than all of them.
That refuses anything needing one, which is wrong but visible -- where assuming
it can do everything would assign work it cannot do and find out on the
machine. And a capability the node reported as ABSENT is not counted: reading
the list without the verdict would let a module onto a machine that said no.

`overlay push` is gone, replaced by `push`, which sends a node its network and
its modules as one declaration. Two commands that overlap is how a mesh ends up
half-configured by whichever was run. The old name answers with where to go,
and answers before opening a database -- needing one would turn a redirect into
a connection error.
2026-08-29 22:16:11 +02:00