Commit Graph
20 Commits
Author SHA1 Message Date
jschoubben 964285f08c The build machine takes work on the bus its credential names, and the work queue has a taker
Two halves of one gap the first build over the new bus met. The machine decided
its bus from a variable its container never received, so the credential the mesh
sealed to it went unread; a credential for the new bus names the bus by scheme and
carries user, password and fingerprint beside the address, and that is enough to
dial it, pinned. And the roles' work queues were raised with no holders, so the
consumer a machine binds to take work was never created: the holders are read
from the catalogue and the handover record, as the resolver reads them.
2026-09-28 01:59:20 +02:00
jschoubben e4e960ec1c A build is work submitted to a role, on both buses
ADR 0121 carried through to working code. `Builders` is the asking side and
`BuildMachine` the taking side, each with an implementation per bus, and the builder
binary and the `build` command now go through them.

On the bus being built, one publish does what two did. The old bus answered the asker
through a reply queue and announced to an events exchange, because two audiences meant
two topologies. Here the outcome is the role's own event: the asker matches it by the
id its request carried, the controller records it, the catalogue places it in the graph.
So a build machine publishes once, needs a reply queue for nothing, and needs a grant
over nobody's inbox — which is what ruled out the alternatives.

The outcome carries the module name now. Only the manifest says what was built, and on
the old bus the separate announcement carried it; with one message for three readers it
belongs in the result. A failed build names none, because it produced no module version
and the catalogue would otherwise place something that was never made.

Checked against a real server: the whole round trip; a third party on the role's event
hearing the same outcome the asker did, which is the claim the decision rests on; work
leaving the queue once settled, so no second machine repeats it; work submitted with no
machine holding the role waiting instead of failing, and being done when one arrives;
and work a machine handed back coming round again.

One thing I got wrong twice now and have written down where it bit: binding to a
consumer must name that consumer's own filter subject, not the narrower subject the
caller cares about. The client compares the two and refuses anything that is not equal,
with "subject does not match consumer".
2026-09-27 16:01:53 +02:00
jschoubben 6c12780abe Describe the bus on its own terms
Comments framed the new bus by what it replaces — a comparison in almost
every explanation, which reads as though NATS were a variant of the old
thing rather than the mesh's nervous system. Removed throughout, and
OverAMQP becomes OverCurrent: the seam's two sides are the bus the mesh
runs on today and the one being built, not two protocols.

What remains is the client library's own package name, which is its name.
2026-09-26 23:51:00 +02:00
jschoubben 2fad32767e The controller's outbound link behind a seam, with both transports
Step 3.4, first half. Every one of these took an *amqp.Channel, so the
transport reached every caller and swapping it meant touching all of them.
The seam turned out to be small — the controller sends exactly two kinds of
message that expect no answer — which is the same measurement that said
this bus could be replaced at all.

Bus is stated in the mesh's words, not a transport's: PublishEvent and
PublishDeclaration. Two implementations, both shipping, because steps 1 to
4 leave every node on AMQP and the NATS one is selected at the rollout.
Both ship is also what makes them comparable: one conformance fixture holds
both to the same envelope, and the NATS one is checked against a real
server reading back from the stream rather than from the code that wrote it.

Still on *amqp.Channel: RequestBuild and Ask, which carry reply-queue
machinery, and the whole consume side — the control loop, enrolment, serve.
2026-09-26 23:47:15 +02:00
jschoubben 6ac9013d6e builder: a clone may offer the forge's credential, through git's own store
A private repository could not be built: the builder clones anonymously,
and had no way to say who it is. It already holds exactly one credential
to exactly the right place — the package-registry binding and its sealed
secret, one gitea user whose password answers npm and git alike — so a
clone now offers that, and nothing new is minted or carried.

Offered, never pushed: the credential is written as a git
credential-store file (0600, in the workspace, never argv) and named
with -c credential.helper, so git itself decides when it applies — only
on an authentication challenge, and only for the URL it was written
for, scheme, host and port included. A public repository clones exactly
as before; a repository on any other host is never shown it. The same
store rides along on an artifact's own context clone, so a private
module with a private context builds too.
2026-09-25 21:47:32 +02:00
jschoubben c3b88b9148 Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben 4b9bc50aad The builder resolves the SDK from the mesh registry, and can publish packages
A new 'package' artifact kind builds a module's own code on a public base image
and publishes it to the mesh's package registry by version (hq ADR 0076) — the
SDK above all, which the toolchain is built from and so cannot be built in the
toolchain. The credential a build needs to resolve or publish packages is
rendered as an .npmrc (basic auth, hq ADR 0048) and given to an image build as a
buildkit secret, never a layer, so a token is not baked into the toolchain image.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 10:27:26 +02:00
jschoubben 2fb700d61b Builder diagnostics go to stderr; stdout is the result alone
The logging added a line to stdout, and the genesis path parses the builder's
stdout as JSON — so the first log line broke the parse with "invalid character
'c'", the c from "[clone]". A build that had worked stopped working because of a
print statement.

The installer's runner captures stdout alone (cmd.Output), and the contract was
already stdout=result, stderr=everything else. The fix is to honour it: every
builder diagnostic — the step log, the per-command echo, the module path's own
lines — goes to stderr. Stdout carries only once.go's result JSON.

And a unit test now fails if any fmt.Print to stdout appears in the two builder
command files, except the three that belong there: the result, --version, and
--help. A guard, because this was invisible until a 20-minute run hit it, and the
same class of mistake should fail in milliseconds next time.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 23:01:04 +02:00
jschoubben 9070d2502c The builder narrates every step, and every command it runs
A build was silent from clone to publish, so a build in progress, one that failed
quietly, and a request that never arrived all looked identical — which cost a long
diagnosis against a running mesh chasing "the handler never fired".

Now: the handler announces a request the instant it lands. Build logs each phase
— clone, commit, manifest, bases, each artifact starting and finishing with what
it produced, resolve, done — through a Log callback that is nil-safe, so the tests
that pass none still build. And the Command runner echoes every command before it
runs, with where and how long it took, because on a hang the last line is exactly
the command it is stuck inside: "git clone waiting on a network that will not
answer" rather than "the builder did nothing".

The unreadable-request path prints to stdout now too, not stderr, so it shows in
docker logs without splitting streams — the split is what hid it.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 22:26:35 +02:00
jschoubben cfe2816495 A module names the module its build stands on, not a copy of it
A fingerprint written into a recipe names one particular copy of the base — the
copy on whichever machine the person typing it was using. On any other mesh that
copy has never existed, so the build stops on its first line with a message
about an image nobody can look up. Three modules in the catalogue were in
exactly that state, and the line each of them replaced was equally dead.

A module now names the module and artifact instead, and the mesh answers with
what it holds. The builder is still a thing that clones, builds and answers: the
answer travels with the question, because only the mesh knows what it has.

A base the mesh has not built is refused before anything is built, naming which
module has to exist first.
2026-09-13 23:53:22 +02:00
jschoubben 603be706fc The builder builds one module and stops, with no broker and no registry
This is how a mesh is raised: the installer carries this program and runs it
once, before anything exists, to produce the control plane from the same
repository and path every later rebuild will use. What raises the mesh is then
the same thing that maintains it, rather than a second mechanism exercised once
per new mesh — which is how often enough to rot.

With nowhere to publish, an image stays in the machine's own runtime and is
named by the digest of its own configuration: the same identity the installer
has always used for the image it carried.
2026-09-13 03:24:10 +02:00
jschoubben ca689073c4 A build says which machine made it, in the mesh's name for that machine
An event whose origin reads a container id names something no other module
can look up. The mesh already knows the answer, and a module's environment
file is a file resource, so ${machine:name} reaches it with no composer change.
2026-09-13 01:05:40 +02:00
jschoubben 3ffddff8ed The builder announces what it built, and what it was built on top of
Answering and announcing are different acts. The reply goes to whoever asked and
is correlated to their request; the announcement says to the whole mesh that a
module now exists at a commit, which is what the catalogue places in the module
graph (novox/hq ADR 0072). A build nobody asked for still has to be announced, or
the graph knows less than the registry does.

What it was built on top of is read out of the build's own inputs rather than
declared, because a declared list drifts from what the code actually uses
(ADR 0009). These are artifact references, which is what a build input names;
resolving them to module-versions is the catalogue's work, since it is what knows
which module-version published which artifact.

Events ride the topic exchange, not the direct one nodes speak over, so the
builder's account is granted both: it must be able to answer and to announce.
The envelope is the sdk's, reproduced exactly — a second shape would be a second
thing for consumers to handle, and they are written against the first.

Announcing is not allowed to fail a build. The work was done and was answered; a
build reported as failed because saying so failed is a lie about it.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-13 00:57:41 +02:00
jschoubben f151de103f Build a module from a repository and a path within it
The builder cloned a repository and read the manifest at its root, which means one
repository per module. Nothing we have is shaped that way, so the builder could be
asked to build nothing that exists (novox/hq ADR 0069).

The path travels the whole way — named when asking, carried in the request, used
to read the manifest and as the context everything is produced from, echoed back
in the result, and recorded as part of where a module came from. Without that last
part the mesh could notice a module was behind its source and then be unable to
rebuild it, which is the worst of both.

A path climbing out of the clone is refused: a machine whose job is building other
people's repositories must not read whatever else is on its disk.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:45:50 +02:00
jschoubben e2916ee3db The pin is compared in the spelling the mesh writes it
The builder hashed the certificate and compared a bare digest against a
fingerprint written as `sha256:` followed by 64 hex characters. It could never
match — and it failed as "this is not the broker this builder was told about",
which is the one thing this check exists to report truthfully. A check that
cries wolf on every correct broker is worse than no check, because the first
thing anybody does is remove it.

The error now prints what was expected beside what arrived, the way the host's
has always done: without both, the message describes a mismatch nobody can
confirm.

And the pin check has its own test, driven against a real TLS handshake — it
accepts the certificate whose fingerprint the mesh wrote and refuses another.
A pin only ever exercised through a live broker is a pin nothing tests.
2026-08-31 01:45:09 +02:00
jschoubben a5b11fd6b2 A build machine is told what to check the broker against
The credential was a URL and nothing else, so the builder verified the broker
the ordinary way — against public roots. A mesh's broker presents a certificate
of the mesh's own, which is in no trust store anywhere, so the connection could
only ever succeed against a broker somebody else vouches for. It failed at TLS
with an error about an unknown authority rather than about a missing pin, and
the container sat there running: up, credential on disk, connected to nothing.

So the sealed credential now carries the URL and the broker's fingerprint —
the same two facts a node's token carries, for the same reason, delivered out
of band relative to the thing being trusted. The builder pins it: the standard
chain check is replaced rather than removed, and what replaces it is stricter,
accepting one certificate instead of every certificate a public authority
would sign.

A file holding only a URL still works, for a builder somebody runs by hand
against a broker with an ordinary certificate.
2026-08-31 00:55:33 +02:00
jschoubben 0262873254 status --json, so a board has something to read
A board reads through interfaces and holds nothing. Everything it needs
is already answered — as text, for people, which is not something a page
can read.

`--json` rather than a serving API, because nothing needs one yet:
whatever serves a board runs the command, and the constraint holds either
way — the board never touches a context's store. An API is the larger
thing and should wait until something asks for it.

Both forms are gathered from the same reads before either says anything,
so they answer the same questions rather than being two implementations
that can drift. That was not true of the first version: the JSON printed
after the text, because the branch was too late.

Four properties, each asserted and each confirmed to fail when removed:

- refused and failed stay distinct all the way out. They are fixed in
  different places, so one word for both sends half a page's readers to
  the wrong one — and how much DID apply is carried, since "three of
  eight" and "none of eight" are different machines
- a machine that never spoke carries no time at all, rather than a zero
  one that any page would format as a date in 1970
- nothing is null. A page distinguishing "no machines are wrong" from
  "this field is missing" has to handle both, and null is the one that
  gets forgotten
- no field is named like a secret. Everything here comes from records
  that hold no readable one, but a shape a page is built against is
  exactly where one would eventually be added for convenience
2026-08-30 20:22:04 +02:00
jschoubben 79d6ade4c8 push --behind, and a builder told where to publish
`status` says which machines are not doing what they were told, and
nothing acted on it: a machine that refused or failed stayed wrong until
somebody ran push again naming it.

`push --behind` sends only to machines whose last report was not a clean
apply. A command rather than a timer, deliberately: a scheduler is then a
scheduler over this, where building the scheduler first would have meant
two paths to one act with nothing to compare them against.

Naming a machine and asking which machines need one are different
requests, so `push <node> --behind` is refused rather than guessed. With
nothing behind it says so, because "nothing needed one" and "this did not
run" must never look the same. A machine failing the same way for six
hours is pushed to anyway and said about — refusing would leave no way to
retry after fixing the cause, and this is a command somebody ran.

Proven in the lab: a machine is broken with a package that does not
exist, `push --behind` names it and not the machine that is fine, the
module is corrected, and the machine recovers without anybody naming it.

And the builder can be told where to publish rather than configured. A
builder that is a module requires an artifact store, and the mesh writes
it the same binding any consumer of any provision gets. A binding with no
address is refused rather than falling back to anything — that would
publish to a store on the wrong machine and be found out much later. The
variable remains for a builder run by a person, which is how it is still
run while being developed.
2026-08-30 19:47:02 +02:00
jschoubben 3195634441 A build machine gets its own credential, scoped to build work
The builder was documented as holding its own broker credential and
nothing else, and nothing issued one — so in practice it used whatever it
was handed, which was the broker's administrative account. A program
documented as holding its own credential and given somebody else's is
worse than one with no story at all.

`builder issue <name>` creates an account that may read the build queue
and write to the mesh exchange. Not a node account: a build machine is
not a node, and a node's queue carries its declarations.

Two faults found by running it, both about the answer path:

- the reply queue was left for the broker to name, and the account was
  scoped to `amq.gen-*` — one broker's convention. The builder built,
  could not answer, and the connection closed. Reply queues are named
  here now, deterministically.
- the answer then went via the DEFAULT exchange, where permission is
  granted per exchange rather than per queue. A builder allowed to use it
  could publish into any node's queue, which is the privilege a build
  machine most obviously should not have. Answers go through the mesh
  exchange, which it already may use, and an asker binds its reply queue
  to the same key and filters by correlation.

Verified against a real broker: a builder cannot consume a node's queue
and cannot publish to the default exchange. That check nearly reported
the opposite — an unconfirmed publish is asynchronous, so the refusal
arrives as a channel close afterwards and a naive test sees success. With
publisher confirms it is immediate. A negative security assertion made
against an asynchronous call is not an assertion.

Redelivery was observed working while fixing this: builders that died
before answering left their work on the queue, and the next builder did
all of it.

Also: the queue and exchange names exist in both `broker` and `link`,
because `link` imports `broker`. A test in an external package keeps them
agreeing — a builder scoped to a queue nothing publishes to takes no work
and says nothing about why.
2026-08-30 10:28:41 +02:00
jschoubben 421fe73dce The mesh builds: a machine takes the work, and the catalogue shows it
A build is work, not state. Everything else the control plane sends a
node is a declaration — this is what you should be — reconciled forever.
A build happens once and is finished. Putting it in a declaration would
mean rebuilding on every reconcile, or a declaration carrying "and I
already did this", which is state about an event rather than about a
machine.

So it travels on its own queue and the answer comes back correlated. One
queue, so several build machines share the work and each request is done
exactly once — which a per-machine routing key would not give.

mesh-builder is the program a build machine runs. Not the control plane,
which must not run commands on a machine; not the host, which would then
need a container runtime and git everywhere to do something almost no
machine will ever do. It holds its own broker credential and nothing
else.

Three properties that are decisions:

- a request is acknowledged only once the answer is away, so a builder
  that dies mid-build leaves the work for another machine rather than
  losing it with nobody ever hearing why
- one build at a time. Five at once against one runtime finishes all five
  slower than it would have finished the first, and the queue is what
  shares work between machines
- a failure is a RESULT. A build that fails silently is
  indistinguishable from a builder that is not running, and those want
  different responses

And `module list` is a catalogue: what exists, at which version, built
from which commit or handed over by hand or shipped with the control
plane, whether it is behind its source, and which machines run it. All of
that was recorded from the first build and none of it was shown, so "is
this current?" could only be answered by reading the database.

Proven against a real broker, registry and store: the mesh asked, a
builder consumed, built, published, answered; the manifest was recorded
with its commit; the source moved and the catalogue said "behind";
rebuilding caught it up with a new digest because the content changed.
2026-08-30 03:46:02 +02:00