Both assertions were wrong and the mesh was right, which the output made
plain: the rule set named its source, dropped by default, restricted the
declared port and omitted the undeclared one.
"From the mesh" resolves to the addresses on the private network — the whole
point — and the assertion was looking for the segment the two machines happen
to share. So the test now reaches the same machine both ways, and asserts the
declared port answers over the private network and does NOT answer off it.
A test with only one path could not tell "open to the mesh" from "open".
And the fingerprint is delivered with a sha256: prefix, which the regex did not
allow.
A builder that cannot connect sits there, and every outward sign — container
up, credential on disk — says it is working. The failure surfaced five minutes
later as nothing consuming the build queue, which names no cause at all.
Two faults in one line of the harness, and the second is the serious one.
Every command was wrapped as `<cmd> 2>&1; echo "__exit=$?"` on a single line,
so any command containing a heredoc broke: the terminator line became
`MARKER 2>&1; echo ...`, matched nothing, and the heredoc swallowed the rest of
the script — the echo with it. `exec 2>&1` on its own first line fixes that: a
heredoc then terminates where it says it does.
And when the marker was gone, `Number("")` is 0, so the missing exit status
read as exit 0. A command whose output was swallowed reported that it worked,
which is the one answer a test harness must never give. It is now a failure,
with whatever was said returned so the reason is visible.
Found because the firewall test's listener is written with a heredoc and never
started, and the test failed on its own setup — which reads exactly like the
firewall working.
Two setup faults, each of which looked like the thing being tested failing.
The builder module was assigned without its artifact ever being built, so
nothing could start — and the build has to happen while the hand-started
builder is still alive. Same chicken-and-egg as the registry, resolved the same
way: the builder that exists builds the one that replaces it.
The firewall test's listeners were squeezed through three levels of shell
quoting and never started, so the test failed on its own setup — which reads
exactly like the firewall working.
Written and loaded are different things, and loaded and enforcing are different
again. The test opens two ports on a machine, declares one of them, and checks
from the other machine that the declared one answers and the undeclared one
does not — then removes the module and checks the port closes with nobody
editing a rule.
The base image gains nftables, read back through `nft --version` like the other
three: a machine that cannot load a rule set applies the mesh's filtering,
reports success and filters nothing, which is the exact fault the derivation
exists to remove.
Two earlier tests were asking for things that are not there. The lab's registry
drops tags when it stocks, so `registry:2` is not served and the mirror test
failed with "not found" — it now uses the pinned digest, which is what a
declaration carries anyway.
went
`exec` waited two minutes always. A build, or anything that waits on
another machine, needs longer — and a caller that cannot say so has to
split the work to fit, which is a test shaped by its harness rather than
by what it is testing.
The scenario also places a build machine when one is given, so anything
in it can ask the mesh to build something. Nothing else here would start
one.
And the mesh-runs-its-own-artifact-store test is not here. It needs a
fourth image so the module has a registry to mirror, and that is caught
behind 04-ISSUES/012 — left as a note saying where it went and why,
rather than silently deleted, because what it asserted is worth
asserting.
Broken with a package that does not exist, so the failure is real and
fixable. The mesh reports it failed; the resources that could be applied
were, because one broken thing no longer blocks the rest; `push --behind`
names that machine and not the one that is fine; the module is corrected;
and the machine recovers with nobody naming it.
And with nothing behind, it says so rather than doing nothing quietly.
Two properties the design claims and neither had been run.
A push to a machine that is switched off must not be lost — a machine is
disconnected as an ordinary situation, not an exception. The queue is
durable and the message persistent, which ought to be enough, but a lost
declaration is silent and "ought to be" is not a property. It waits: the
machine's host is stopped, the push happens, nothing changes on the
machine, and when it listens again it applies what it missed with no
second push and nobody saying anything.
Getting there found a real fault, now fixed in mesh-host and recorded as
04-ISSUES/011: the machine stopped at the first failing resource, so one
broken module blocked every module after it for ever. The evidence was
the broker's queues being EMPTY — the declaration had been delivered and
read.
And removal: two modules assigned, one unassigned, and the machine loses
exactly that one's file while keeping the other's — and keeps the store,
broker and control plane it raised from its own bundle, which the mesh
never declared and must never remove.
Two of my own traps recorded in the test, because both cost real time:
`pkill -f` matches the shell running it, which kills the connection
carrying the command and hangs the caller for ever; and a test that
depends on state another test left behind fails for a reason that has
nothing to do with what it claims.
The status path was demonstrated with inserted rows, which proves the
query and not the path. This sends a real machine something it will
genuinely fail at — a package that does not exist — and asks the mesh
afterwards.
A failure of the ordinary kind: the host tries, the package manager says
no, some of the declaration is applied and some is not. That is the
situation `status` exists to distinguish from a machine that refused
everything, and the test asserts the distinction survives the whole way:
the machine is listed as failed rather than refused, the failing resource
is named in the host's own words, and the machine that did as it was told
is not implicated.
Everything before this proved a part. This proves the parts meet, which
the project keeps saying cannot be checked any other way.
A bare machine applies the substrate bundle and becomes a mesh — store,
schemas, broker with a certificate it generated itself, control plane
serving. Both machines then join it with nothing but a token. A database
is declared on one and an application on the other, and after a push:
- both ends hold the SAME password, or nothing could authenticate
- it is mode 0600 on the machine that uses it
- it appears in neither machine's stored declaration, neither machine's
reported state, nor the control plane's database — so it was not
readable by the broker that carried it or the mesh that sent it
- the consumer is also told where its database is, by a name the mesh
wrote into that machine's hosts file
The bundle's image references are rewritten to the ones this scenario's
registry serves. A digest belongs to whatever registry serves it, so a
committed bundle names a registry that is not this one — rewriting is
what makes it applicable rather than a placeholder to tidy away.
Two faults found getting here, both fixed in mesh-host: `apply` could not
read a file the bundle could, and the token did not say what the mesh
calls the machine.