Two properties the design claims and neither had been run.
A push to a machine that is switched off must not be lost — a machine is
disconnected as an ordinary situation, not an exception. The queue is
durable and the message persistent, which ought to be enough, but a lost
declaration is silent and "ought to be" is not a property. It waits: the
machine's host is stopped, the push happens, nothing changes on the
machine, and when it listens again it applies what it missed with no
second push and nobody saying anything.
Getting there found a real fault, now fixed in mesh-host and recorded as
04-ISSUES/011: the machine stopped at the first failing resource, so one
broken module blocked every module after it for ever. The evidence was
the broker's queues being EMPTY — the declaration had been delivered and
read.
And removal: two modules assigned, one unassigned, and the machine loses
exactly that one's file while keeping the other's — and keeps the store,
broker and control plane it raised from its own bundle, which the mesh
never declared and must never remove.
Two of my own traps recorded in the test, because both cost real time:
`pkill -f` matches the shell running it, which kills the connection
carrying the command and hangs the caller for ever; and a test that
depends on state another test left behind fails for a reason that has
nothing to do with what it claims.
The status path was demonstrated with inserted rows, which proves the
query and not the path. This sends a real machine something it will
genuinely fail at — a package that does not exist — and asks the mesh
afterwards.
A failure of the ordinary kind: the host tries, the package manager says
no, some of the declaration is applied and some is not. That is the
situation `status` exists to distinguish from a machine that refused
everything, and the test asserts the distinction survives the whole way:
the machine is listed as failed rather than refused, the failing resource
is named in the host's own words, and the machine that did as it was told
is not implicated.
Everything before this proved a part. This proves the parts meet, which
the project keeps saying cannot be checked any other way.
A bare machine applies the substrate bundle and becomes a mesh — store,
schemas, broker with a certificate it generated itself, control plane
serving. Both machines then join it with nothing but a token. A database
is declared on one and an application on the other, and after a push:
- both ends hold the SAME password, or nothing could authenticate
- it is mode 0600 on the machine that uses it
- it appears in neither machine's stored declaration, neither machine's
reported state, nor the control plane's database — so it was not
readable by the broker that carried it or the mesh that sent it
- the consumer is also told where its database is, by a name the mesh
wrote into that machine's hosts file
The bundle's image references are rewritten to the ones this scenario's
registry serves. A digest belongs to whatever registry serves it, so a
committed bundle names a registry that is not this one — rewriting is
what makes it applicable rather than a placeholder to tidy away.
Two faults found getting here, both fixed in mesh-host: `apply` could not
read a file the bundle could, and the token did not say what the mesh
calls the machine.