Commit Graph
34 Commits
Author SHA1 Message Date
jschoubben a83dd3d00a A module says own-secrets, not needs 2026-08-31 13:47:21 +02:00
jschoubben a83f6b7f8e Assign the resolver configuration module that suits the machine
These machines run systemd-resolved, so resolv-conf would fight it over the
file. Both claim the-resolver-configuration so that assigning the wrong one is
refused rather than fought over — and the test was picking the wrong one.
2026-08-31 13:39:58 +02:00
jschoubben a6480bbd88 Say what holds the address, rather than that it is held
dnsmasq cannot bind 127.0.0.54 and the module's own comment claims that address
is free. Which of those is wrong is the question, and naming the holder answers
it — assuming would be issue 012's mistake, where two things changed and the
plausible one was blamed.
2026-08-31 13:24:40 +02:00
jschoubben 89b6dd6080 Print why the resolver did not start, instead of that it did not
`systemctl is-active` exits non-zero for a unit that failed, so `must` threw
before the assertion carrying every diagnostic — and the run said only
"failed". The journal, the config, what the mesh wrote and resolv.conf are all
things the next run should not have to be re-run to see.
2026-08-31 13:11:02 +02:00
jschoubben 20690964f1 Prove a service is reached by a name under the machine it runs on
Through the path an application actually takes — nsswitch, files, then DNS —
because the resolv.conf module is half of what is being tested and only that
path goes through it. Asking a server directly would prove less.

Both machines resolve, from their own copy: a mesh where one machine answers
for all of them stops resolving when that machine does, which is the
arrangement this design refuses everywhere else.

The manifests are read from mesh-control's examples rather than written here,
so what is proven is what ships. And dnsmasq joins the base image, read back
through --version like the others: a machine that cannot answer names applies
the resolver data, reports success, and resolves nothing.
2026-08-31 12:55:59 +02:00
jschoubben 622b414e4e Take the resolver off before expecting a machine to leave the network
mesh-resolver requires name resolution, which requires the network — so
unassigning the domain module alone leaves the machine on the network, pulled
back by its own requirement. The mesh was right and the test was wrong.

Which is the requirement graph doing its job: a module cannot quietly lose
something it depends on because somebody removed the thing that first brought
it in.
2026-08-31 11:53:16 +02:00
jschoubben cce39a6ba3 Prove every name under a machine resolves to that machine
The mesh's half: the data is right, complete on every machine, and agrees with
the hosts file — two accounts of where a machine is, disagreeing, would be
worse than either alone, and this is the one place they could drift because
they are generated separately.

And it follows the machines: a node that leaves the private network stops being
answered for, because a wildcard pointing at nothing resolves and then hangs,
where an unresolvable name fails at once and says which name it was.
2026-08-31 11:40:19 +02:00
jschoubben 8bc5f498cd Put the workaround back where the container is not the mesh's
The mesh gives its names to the containers it declares. This one is started by
the test with `docker run` — nothing declared it, so nothing configured it, and
removing the workaround here was claiming a reach the change does not have.

The boundary is the right one: a container somebody runs by hand is not the
mesh's to configure. Reaching into every container on a machine, declared or
not, is what a resolver in resolv.conf would be for — and that remains the
case for wanting one.

The proof that names work inside containers is its own test, against a
container the mesh declared, and it passes.
2026-08-31 11:28:38 +02:00
jschoubben 4726a51986 Reach the other machine by name from inside a container
The rotation test resolved the address on the machine and passed it in, because
the name failed inside the container. That workaround is gone, and a test that
checks the name from inside a container is added — on the machine it has always
worked, which is what made this easy to miss.
2026-08-31 11:13:34 +02:00
jschoubben 0af80ddf4f Name the module that opens the hub's port, not the domain it answers
`networking` is the requirement a module offers; `mesh-wireguard` is the
module, and what caused a rule is the module itself. The rule set was right and
the assertion was looking for the wrong name.
2026-08-31 10:58:38 +02:00
jschoubben 9ab73d6cdb Prove the hub can be filtered without severing the mesh
The failure guarded against is not subtle and is very hard to recover from: a
rule set that closes the hub's own port takes the private network down, and the
mesh's way of fixing anything is to send a declaration over it.

So the assertion that matters is not the rule file — it is that a declaration
still reaches the other machine afterwards, and that the other machine still
reaches the hub. A rule file that looks right and a mesh that has stopped are
exactly what this is for.
2026-08-31 10:07:09 +02:00
jschoubben 5736cca9f6 Break a machine with a unit that does not exist
A file in a missing directory is not impossible: the host creates the parents,
which is correct and meant the first version of this test broke nothing at all
— and then reported that the board could not name a failure that never
happened. A unit that does not exist fails immediately and in the host's own
words, which is also what the board is being asked to show.
2026-08-31 05:01:53 +02:00
jschoubben eb0c8a6d98 Prove the board says what a person opened it for
A machine is given a declaration it cannot apply, and the board names it,
says 'failed' rather than 'error', and shows the host's own words about what it
could not do — a board that said only 'failed' would send a person to ask the
thing they opened the board to avoid asking.

And it agrees with the command, from the same read: two answers to 'which
machine is broken' would be worse than either alone. Reading it changes
nothing.
2026-08-31 04:50:22 +02:00
jschoubben 4d5b190db8 Wait for the route to be withdrawn rather than sleeping through it
This assertion passed twice and failed once on nothing but timing, which is the
worst kind of green: it says the mechanism works when what it measured was the
clock.

`status` cannot stand in for the wait either. A machine that has not applied
yet is not a machine that failed — "not yet" and "never" look identical there,
and only one of them is worth failing over. So it waits for the thing itself,
and says on failure that the machine did apply, which is what separates "the
mesh still thinks this contributes" from "nothing was sent".
2026-08-31 03:45:31 +02:00
jschoubben eaa7cdea79 Check the provider applied before reading what its file says
The mesh withdrawing a route and the machine acting on it are different things,
and a test that reads the file without checking the second reports the first
wrongly whenever the machine is behind for any unrelated reason. On failure it
now also prints what the mesh would send now, which is what separates 'the mesh
still thinks this contributes' from 'the machine never applied'.
2026-08-31 03:34:51 +02:00
jschoubben 56de227673 Connect by address, and ask grep whether rather than how many
A container does not inherit its host's /etc/hosts, so a name the mesh wrote
there resolves for the machine and not for anything it runs. It fails as "could
not translate host name", which reads like a mesh that never wrote the name —
so the address is resolved on the machine and the container is given that.

And `grep -c` prints 0 and exits non-zero when it finds nothing, so the obvious
`|| echo 0` prints a second one and the count is never what it looks like. The
question was always whether, not how many.

The rotation check now also reports what psql said, not only what the
provisioner said: the failure was on the client side and the diagnostics were
all from the server.
2026-08-31 03:22:24 +02:00
jschoubben 7669d1c373 Log in as the role the provisioner made, and record the licences first
Two setup faults, each of which read as the mesh failing.

The provisioner names a role after the machine and a database after what the
module asked for. The rotation test logged in as the module's name into the
wrong database, so a provisioner that had done its job exactly looked like one
that had not.

And the licence test assigned before recording any licence, so the refusal it
got was "nothing provides model-access" — correct, and a different refusal from
the one being tested. Assigning is also the earliest point a person meets it,
so that is where it is now checked.
2026-08-31 03:07:31 +02:00
jschoubben 4ee769dda0 Prove a commit reaches a machine already running the old one
Commit, build, catalogue, push — and the machine ends up running what the
source says. With the two halves that make the answer trustworthy: it is still
running the old one until it is told, because the mesh changing its mind is not
a machine acting on it; and it stops being reported as behind once it has
caught up, because a status that says "behind" for ever is one nobody reads.
2026-08-31 02:55:10 +02:00
jschoubben d56a0c8fef Prove a licence is accepted, sealed, and unreadable by the mesh
Refused until somebody says which, with both candidates and the command named.
Refused again while it has no key. Then the key is given on standard input, the
public half arrives saying it came from a record rather than a machine, the key
itself arrives readable only by that machine — and it is nowhere in the control
plane's own database, nor in what crossed the broker.

And the rotation test asked for grants without receives, so the provisioner
found a directory of unexplained secrets and said nothing had been granted:
true, and indistinguishable from a credential never delivered.
2026-08-31 02:52:26 +02:00
jschoubben 47d990b33a Prove a route reaches the workload, and does not outlive it
The request goes to the name, across the private network, and returns the
workload's own answer. Then the module is unassigned and the same request must
stop working — a stale public name pointing at nothing fails more visibly than
a stale grant.

The workload declares its port as well as its route, because they are different
questions and the earlier test leaves this machine filtering: a module that
asked for a route and not for the port would be unreachable by the proxy it
just asked for.
2026-08-31 02:43:19 +02:00
jschoubben 21a1e85d32 Prove rotation against a real database, with a real login
Two ends holding a matching string proves they agree, not that either is right.
So the check is three logins over the private network from the consumer's own
machine: the delivered credential works, the rotated one works, and the one
that was rotated away does not. Without the last, the test passes against a
provider that added a password without replacing one.

Not over loopback: pg_hba trusts anything there, and a deliberately wrong
password returned a row for a whole afternoon once.
2026-08-31 02:39:00 +02:00
jschoubben f5619b02d6 A builder that is a module cannot see the machine's filesystem
It runs in a container, so a path like /root exists for the machine and not for
it. The first build in this test works because the hand-started builder runs on
the host; the second is done by the module, and asked it to clone a path it has
no way to reach.

A real module is cloned from the forge over a URL. The lab has no forge, so the
repository goes in the directory the module already mounts — the same fact
wearing different clothes.
2026-08-31 01:53:22 +02:00
jschoubben 83727099eb The firewall module ships the unit that loads its rules
The distribution's nftables.service is Type=oneshot with no RemainAfterExit: it
loads the rules and goes inactive. A host asked for a service that is "running"
then reports, quite correctly, that it is stopped — every packet filtered as
declared, and the machine marked as not doing what it was told.

There is no state in the vocabulary for "ran and exited having done its job",
so a module that needs one brings a unit that stays. That is also the right
shape: how a machine enforces rules is a fact about the machine, and the mesh
has no business depending on what a distribution happens to package.

And when the builder's build times out, dump the builder's own account of
itself. "Nothing consumed the queue" names no cause and is the same sentence
whether the credential was refused, the queue was never declared, or the
process died three seconds in.
2026-08-31 01:20:46 +02:00
jschoubben 6bd7833ae9 Test that "from the mesh" is not a synonym for "open"
Both assertions were wrong and the mesh was right, which the output made
plain: the rule set named its source, dropped by default, restricted the
declared port and omitted the undeclared one.

"From the mesh" resolves to the addresses on the private network — the whole
point — and the assertion was looking for the segment the two machines happen
to share. So the test now reaches the same machine both ways, and asserts the
declared port answers over the private network and does NOT answer off it.
A test with only one path could not tell "open to the mesh" from "open".

And the fingerprint is delivered with a sha256: prefix, which the regex did not
allow.
2026-08-31 01:07:06 +02:00
jschoubben 35000a236c Assert the builder can reach the broker, not merely that it is running
A builder that cannot connect sits there, and every outward sign — container
up, credential on disk — says it is working. The failure surfaced five minutes
later as nothing consuming the build queue, which names no cause at all.
2026-08-31 00:55:33 +02:00
jschoubben 9711a90bde A command with no marker is a failure, not a success
Two faults in one line of the harness, and the second is the serious one.

Every command was wrapped as `<cmd> 2>&1; echo "__exit=$?"` on a single line,
so any command containing a heredoc broke: the terminator line became
`MARKER 2>&1; echo ...`, matched nothing, and the heredoc swallowed the rest of
the script — the echo with it. `exec 2>&1` on its own first line fixes that: a
heredoc then terminates where it says it does.

And when the marker was gone, `Number("")` is 0, so the missing exit status
read as exit 0. A command whose output was swallowed reported that it worked,
which is the one answer a test harness must never give. It is now a failure,
with whatever was said returned so the reason is visible.

Found because the firewall test's listener is written with a heredoc and never
started, and the test failed on its own setup — which reads exactly like the
firewall working.
2026-08-31 00:48:00 +02:00
jschoubben 44f088a5b6 Assert the build was recorded, not that the word appears
`builds` says "nothing has been built yet" when there is nothing, and the
assertion was matching on a word that sentence contains.
2026-08-31 00:46:09 +02:00
jschoubben 0100c39845 Build the builder before replacing the hand-started one, and listen from a file
Two setup faults, each of which looked like the thing being tested failing.

The builder module was assigned without its artifact ever being built, so
nothing could start — and the build has to happen while the hand-started
builder is still alive. Same chicken-and-egg as the registry, resolved the same
way: the builder that exists builds the one that replaces it.

The firewall test's listeners were squeezed through three levels of shell
quoting and never started, so the test failed on its own setup — which reads
exactly like the firewall working.
2026-08-31 00:41:24 +02:00
jschoubben 29cdaa4de3 Prove a machine filters what it was told to and nothing else
Written and loaded are different things, and loaded and enforcing are different
again. The test opens two ports on a machine, declares one of them, and checks
from the other machine that the declared one answers and the undeclared one
does not — then removes the module and checks the port closes with nobody
editing a rule.

The base image gains nftables, read back through `nft --version` like the other
three: a machine that cannot load a rule set applies the mesh's filtering,
reports success and filters nothing, which is the exact fault the derivation
exists to remove.

Two earlier tests were asking for things that are not there. The lab's registry
drops tags when it stocks, so `registry:2` is not served and the mirror test
failed with "not found" — it now uses the pinned digest, which is what a
declaration carries anyway.
2026-08-31 00:25:19 +02:00
jschoubben 2f81701a13 Let a caller say how long to wait, and note where the artifact-store test
went

`exec` waited two minutes always. A build, or anything that waits on
another machine, needs longer — and a caller that cannot say so has to
split the work to fit, which is a test shaped by its harness rather than
by what it is testing.

The scenario also places a build machine when one is given, so anything
in it can ask the mesh to build something. Nothing else here would start
one.

And the mesh-runs-its-own-artifact-store test is not here. It needs a
fourth image so the module has a registry to mirror, and that is caught
behind 04-ISSUES/012 — left as a note saying where it went and why,
rather than silently deleted, because what it asserted is worth
asserting.
2026-08-30 20:41:10 +02:00
jschoubben c98ee7a82d A machine that fell behind catches up without being named
Broken with a package that does not exist, so the failure is real and
fixable. The mesh reports it failed; the resources that could be applied
were, because one broken thing no longer blocks the rest; `push --behind`
names that machine and not the one that is fine; the module is corrected;
and the machine recovers with nobody naming it.

And with nothing behind, it says so rather than doing nothing quietly.
2026-08-30 19:47:03 +02:00
jschoubben e0793c17f7 A declaration waits, and unassigning takes away exactly what it should
Two properties the design claims and neither had been run.

A push to a machine that is switched off must not be lost — a machine is
disconnected as an ordinary situation, not an exception. The queue is
durable and the message persistent, which ought to be enough, but a lost
declaration is silent and "ought to be" is not a property. It waits: the
machine's host is stopped, the push happens, nothing changes on the
machine, and when it listens again it applies what it missed with no
second push and nobody saying anything.

Getting there found a real fault, now fixed in mesh-host and recorded as
04-ISSUES/011: the machine stopped at the first failing resource, so one
broken module blocked every module after it for ever. The evidence was
the broker's queues being EMPTY — the declaration had been delivered and
read.

And removal: two modules assigned, one unassigned, and the machine loses
exactly that one's file while keeping the other's — and keeps the store,
broker and control plane it raised from its own bundle, which the mesh
never declared and must never remove.

Two of my own traps recorded in the test, because both cost real time:
`pkill -f` matches the shell running it, which kills the connection
carrying the command and hangs the caller for ever; and a test that
depends on state another test left behind fails for a reason that has
nothing to do with what it claims.
2026-08-30 19:39:46 +02:00
jschoubben 230074665f A machine that cannot do what it was told, and the mesh saying so
The status path was demonstrated with inserted rows, which proves the
query and not the path. This sends a real machine something it will
genuinely fail at — a package that does not exist — and asks the mesh
afterwards.

A failure of the ordinary kind: the host tries, the package manager says
no, some of the declaration is applied and some is not. That is the
situation `status` exists to distinguish from a machine that refused
everything, and the test asserts the distinction survives the whole way:
the machine is listed as failed rather than refused, the failing resource
is named in the host's own words, and the machine that did as it was told
is not implicated.
2026-08-30 18:15:14 +02:00
jschoubben 99b3444b19 Two machines, one mesh, a credential neither end had to be told twice
Everything before this proved a part. This proves the parts meet, which
the project keeps saying cannot be checked any other way.

A bare machine applies the substrate bundle and becomes a mesh — store,
schemas, broker with a certificate it generated itself, control plane
serving. Both machines then join it with nothing but a token. A database
is declared on one and an application on the other, and after a push:

- both ends hold the SAME password, or nothing could authenticate
- it is mode 0600 on the machine that uses it
- it appears in neither machine's stored declaration, neither machine's
  reported state, nor the control plane's database — so it was not
  readable by the broker that carried it or the mesh that sent it
- the consumer is also told where its database is, by a name the mesh
  wrote into that machine's hosts file

The bundle's image references are rewritten to the ones this scenario's
registry serves. A digest belongs to whatever registry serves it, so a
committed bundle names a registry that is not this one — rewriting is
what makes it applicable rather than a placeholder to tidy away.

Two faults found getting here, both fixed in mesh-host: `apply` could not
read a file the bundle could, and the token did not say what the mesh
calls the machine.
2026-08-30 03:05:07 +02:00