Commit Graph
4 Commits
Author SHA1 Message Date
jschoubben 0262873254 status --json, so a board has something to read
A board reads through interfaces and holds nothing. Everything it needs
is already answered — as text, for people, which is not something a page
can read.

`--json` rather than a serving API, because nothing needs one yet:
whatever serves a board runs the command, and the constraint holds either
way — the board never touches a context's store. An API is the larger
thing and should wait until something asks for it.

Both forms are gathered from the same reads before either says anything,
so they answer the same questions rather than being two implementations
that can drift. That was not true of the first version: the JSON printed
after the text, because the branch was too late.

Four properties, each asserted and each confirmed to fail when removed:

- refused and failed stay distinct all the way out. They are fixed in
  different places, so one word for both sends half a page's readers to
  the wrong one — and how much DID apply is carried, since "three of
  eight" and "none of eight" are different machines
- a machine that never spoke carries no time at all, rather than a zero
  one that any page would format as a date in 1970
- nothing is null. A page distinguishing "no machines are wrong" from
  "this field is missing" has to handle both, and null is the one that
  gets forgotten
- no field is named like a secret. Everything here comes from records
  that hold no readable one, but a shape a page is built against is
  exactly where one would eventually be added for convenience
2026-08-30 20:22:04 +02:00
jschoubben 79d6ade4c8 push --behind, and a builder told where to publish
`status` says which machines are not doing what they were told, and
nothing acted on it: a machine that refused or failed stayed wrong until
somebody ran push again naming it.

`push --behind` sends only to machines whose last report was not a clean
apply. A command rather than a timer, deliberately: a scheduler is then a
scheduler over this, where building the scheduler first would have meant
two paths to one act with nothing to compare them against.

Naming a machine and asking which machines need one are different
requests, so `push <node> --behind` is refused rather than guessed. With
nothing behind it says so, because "nothing needed one" and "this did not
run" must never look the same. A machine failing the same way for six
hours is pushed to anyway and said about — refusing would leave no way to
retry after fixing the cause, and this is a command somebody ran.

Proven in the lab: a machine is broken with a package that does not
exist, `push --behind` names it and not the machine that is fine, the
module is corrected, and the machine recovers without anybody naming it.

And the builder can be told where to publish rather than configured. A
builder that is a module requires an artifact store, and the mesh writes
it the same binding any consumer of any provision gets. A binding with no
address is refused rather than falling back to anything — that would
publish to a store on the wrong machine and be found out much later. The
variable remains for a builder run by a person, which is how it is still
run while being developed.
2026-08-30 19:47:02 +02:00
jschoubben 3195634441 A build machine gets its own credential, scoped to build work
The builder was documented as holding its own broker credential and
nothing else, and nothing issued one — so in practice it used whatever it
was handed, which was the broker's administrative account. A program
documented as holding its own credential and given somebody else's is
worse than one with no story at all.

`builder issue <name>` creates an account that may read the build queue
and write to the mesh exchange. Not a node account: a build machine is
not a node, and a node's queue carries its declarations.

Two faults found by running it, both about the answer path:

- the reply queue was left for the broker to name, and the account was
  scoped to `amq.gen-*` — one broker's convention. The builder built,
  could not answer, and the connection closed. Reply queues are named
  here now, deterministically.
- the answer then went via the DEFAULT exchange, where permission is
  granted per exchange rather than per queue. A builder allowed to use it
  could publish into any node's queue, which is the privilege a build
  machine most obviously should not have. Answers go through the mesh
  exchange, which it already may use, and an asker binds its reply queue
  to the same key and filters by correlation.

Verified against a real broker: a builder cannot consume a node's queue
and cannot publish to the default exchange. That check nearly reported
the opposite — an unconfirmed publish is asynchronous, so the refusal
arrives as a channel close afterwards and a naive test sees success. With
publisher confirms it is immediate. A negative security assertion made
against an asynchronous call is not an assertion.

Redelivery was observed working while fixing this: builders that died
before answering left their work on the queue, and the next builder did
all of it.

Also: the queue and exchange names exist in both `broker` and `link`,
because `link` imports `broker`. A test in an external package keeps them
agreeing — a builder scoped to a queue nothing publishes to takes no work
and says nothing about why.
2026-08-30 10:28:41 +02:00
jschoubben 421fe73dce The mesh builds: a machine takes the work, and the catalogue shows it
A build is work, not state. Everything else the control plane sends a
node is a declaration — this is what you should be — reconciled forever.
A build happens once and is finished. Putting it in a declaration would
mean rebuilding on every reconcile, or a declaration carrying "and I
already did this", which is state about an event rather than about a
machine.

So it travels on its own queue and the answer comes back correlated. One
queue, so several build machines share the work and each request is done
exactly once — which a per-machine routing key would not give.

mesh-builder is the program a build machine runs. Not the control plane,
which must not run commands on a machine; not the host, which would then
need a container runtime and git everywhere to do something almost no
machine will ever do. It holds its own broker credential and nothing
else.

Three properties that are decisions:

- a request is acknowledged only once the answer is away, so a builder
  that dies mid-build leaves the work for another machine rather than
  losing it with nobody ever hearing why
- one build at a time. Five at once against one runtime finishes all five
  slower than it would have finished the first, and the queue is what
  shares work between machines
- a failure is a RESULT. A build that fails silently is
  indistinguishable from a builder that is not running, and those want
  different responses

And `module list` is a catalogue: what exists, at which version, built
from which commit or handed over by hand or shipped with the control
plane, whether it is behind its source, and which machines run it. All of
that was recorded from the first build and none of it was shown, so "is
this current?" could only be answered by reading the database.

Proven against a real broker, registry and store: the mesh asked, a
builder consumed, built, published, answered; the manifest was recorded
with its commit; the source moved and the catalogue said "behind";
rebuilding caught it up with a new digest because the content changed.
2026-08-30 03:46:02 +02:00