Commit Graph
5 Commits
Author SHA1 Message Date
jschoubben 2f81701a13 Let a caller say how long to wait, and note where the artifact-store test
went

`exec` waited two minutes always. A build, or anything that waits on
another machine, needs longer — and a caller that cannot say so has to
split the work to fit, which is a test shaped by its harness rather than
by what it is testing.

The scenario also places a build machine when one is given, so anything
in it can ask the mesh to build something. Nothing else here would start
one.

And the mesh-runs-its-own-artifact-store test is not here. It needs a
fourth image so the module has a registry to mirror, and that is caught
behind 04-ISSUES/012 — left as a note saying where it went and why,
rather than silently deleted, because what it asserted is worth
asserting.
2026-08-30 20:41:10 +02:00
jschoubben c98ee7a82d A machine that fell behind catches up without being named
Broken with a package that does not exist, so the failure is real and
fixable. The mesh reports it failed; the resources that could be applied
were, because one broken thing no longer blocks the rest; `push --behind`
names that machine and not the one that is fine; the module is corrected;
and the machine recovers with nobody naming it.

And with nothing behind, it says so rather than doing nothing quietly.
2026-08-30 19:47:03 +02:00
jschoubben e0793c17f7 A declaration waits, and unassigning takes away exactly what it should
Two properties the design claims and neither had been run.

A push to a machine that is switched off must not be lost — a machine is
disconnected as an ordinary situation, not an exception. The queue is
durable and the message persistent, which ought to be enough, but a lost
declaration is silent and "ought to be" is not a property. It waits: the
machine's host is stopped, the push happens, nothing changes on the
machine, and when it listens again it applies what it missed with no
second push and nobody saying anything.

Getting there found a real fault, now fixed in mesh-host and recorded as
04-ISSUES/011: the machine stopped at the first failing resource, so one
broken module blocked every module after it for ever. The evidence was
the broker's queues being EMPTY — the declaration had been delivered and
read.

And removal: two modules assigned, one unassigned, and the machine loses
exactly that one's file while keeping the other's — and keeps the store,
broker and control plane it raised from its own bundle, which the mesh
never declared and must never remove.

Two of my own traps recorded in the test, because both cost real time:
`pkill -f` matches the shell running it, which kills the connection
carrying the command and hangs the caller for ever; and a test that
depends on state another test left behind fails for a reason that has
nothing to do with what it claims.
2026-08-30 19:39:46 +02:00
jschoubben 230074665f A machine that cannot do what it was told, and the mesh saying so
The status path was demonstrated with inserted rows, which proves the
query and not the path. This sends a real machine something it will
genuinely fail at — a package that does not exist — and asks the mesh
afterwards.

A failure of the ordinary kind: the host tries, the package manager says
no, some of the declaration is applied and some is not. That is the
situation `status` exists to distinguish from a machine that refused
everything, and the test asserts the distinction survives the whole way:
the machine is listed as failed rather than refused, the failing resource
is named in the host's own words, and the machine that did as it was told
is not implicated.
2026-08-30 18:15:14 +02:00
jschoubben 99b3444b19 Two machines, one mesh, a credential neither end had to be told twice
Everything before this proved a part. This proves the parts meet, which
the project keeps saying cannot be checked any other way.

A bare machine applies the substrate bundle and becomes a mesh — store,
schemas, broker with a certificate it generated itself, control plane
serving. Both machines then join it with nothing but a token. A database
is declared on one and an application on the other, and after a push:

- both ends hold the SAME password, or nothing could authenticate
- it is mode 0600 on the machine that uses it
- it appears in neither machine's stored declaration, neither machine's
  reported state, nor the control plane's database — so it was not
  readable by the broker that carried it or the mesh that sent it
- the consumer is also told where its database is, by a name the mesh
  wrote into that machine's hosts file

The bundle's image references are rewritten to the ones this scenario's
registry serves. A digest belongs to whatever registry serves it, so a
committed bundle names a registry that is not this one — rewriting is
what makes it applicable rather than a placeholder to tidy away.

Two faults found getting here, both fixed in mesh-host: `apply` could not
read a file the bundle could, and the token did not say what the mesh
calls the machine.
2026-08-30 03:05:07 +02:00