Commit Graph
16 Commits
Author SHA1 Message Date
jschoubben 44f088a5b6 Assert the build was recorded, not that the word appears
`builds` says "nothing has been built yet" when there is nothing, and the
assertion was matching on a word that sentence contains.
2026-08-31 00:46:09 +02:00
jschoubben 0100c39845 Build the builder before replacing the hand-started one, and listen from a file
Two setup faults, each of which looked like the thing being tested failing.

The builder module was assigned without its artifact ever being built, so
nothing could start — and the build has to happen while the hand-started
builder is still alive. Same chicken-and-egg as the registry, resolved the same
way: the builder that exists builds the one that replaces it.

The firewall test's listeners were squeezed through three levels of shell
quoting and never started, so the test failed on its own setup — which reads
exactly like the firewall working.
2026-08-31 00:41:24 +02:00
jschoubben 29cdaa4de3 Prove a machine filters what it was told to and nothing else
Written and loaded are different things, and loaded and enforcing are different
again. The test opens two ports on a machine, declares one of them, and checks
from the other machine that the declared one answers and the undeclared one
does not — then removes the module and checks the port closes with nobody
editing a rule.

The base image gains nftables, read back through `nft --version` like the other
three: a machine that cannot load a rule set applies the mesh's filtering,
reports success and filters nothing, which is the exact fault the derivation
exists to remove.

Two earlier tests were asking for things that are not there. The lab's registry
drops tags when it stocks, so `registry:2` is not served and the mirror test
failed with "not found" — it now uses the pinned digest, which is what a
declaration carries anyway.
2026-08-31 00:25:19 +02:00
jschoubben 2f81701a13 Let a caller say how long to wait, and note where the artifact-store test
went

`exec` waited two minutes always. A build, or anything that waits on
another machine, needs longer — and a caller that cannot say so has to
split the work to fit, which is a test shaped by its harness rather than
by what it is testing.

The scenario also places a build machine when one is given, so anything
in it can ask the mesh to build something. Nothing else here would start
one.

And the mesh-runs-its-own-artifact-store test is not here. It needs a
fourth image so the module has a registry to mirror, and that is caught
behind 04-ISSUES/012 — left as a note saying where it went and why,
rather than silently deleted, because what it asserted is worth
asserting.
2026-08-30 20:41:10 +02:00
jschoubben c98ee7a82d A machine that fell behind catches up without being named
Broken with a package that does not exist, so the failure is real and
fixable. The mesh reports it failed; the resources that could be applied
were, because one broken thing no longer blocks the rest; `push --behind`
names that machine and not the one that is fine; the module is corrected;
and the machine recovers with nobody naming it.

And with nothing behind, it says so rather than doing nothing quietly.
2026-08-30 19:47:03 +02:00
jschoubben e0793c17f7 A declaration waits, and unassigning takes away exactly what it should
Two properties the design claims and neither had been run.

A push to a machine that is switched off must not be lost — a machine is
disconnected as an ordinary situation, not an exception. The queue is
durable and the message persistent, which ought to be enough, but a lost
declaration is silent and "ought to be" is not a property. It waits: the
machine's host is stopped, the push happens, nothing changes on the
machine, and when it listens again it applies what it missed with no
second push and nobody saying anything.

Getting there found a real fault, now fixed in mesh-host and recorded as
04-ISSUES/011: the machine stopped at the first failing resource, so one
broken module blocked every module after it for ever. The evidence was
the broker's queues being EMPTY — the declaration had been delivered and
read.

And removal: two modules assigned, one unassigned, and the machine loses
exactly that one's file while keeping the other's — and keeps the store,
broker and control plane it raised from its own bundle, which the mesh
never declared and must never remove.

Two of my own traps recorded in the test, because both cost real time:
`pkill -f` matches the shell running it, which kills the connection
carrying the command and hangs the caller for ever; and a test that
depends on state another test left behind fails for a reason that has
nothing to do with what it claims.
2026-08-30 19:39:46 +02:00
jschoubben 230074665f A machine that cannot do what it was told, and the mesh saying so
The status path was demonstrated with inserted rows, which proves the
query and not the path. This sends a real machine something it will
genuinely fail at — a package that does not exist — and asks the mesh
afterwards.

A failure of the ordinary kind: the host tries, the package manager says
no, some of the declaration is applied and some is not. That is the
situation `status` exists to distinguish from a machine that refused
everything, and the test asserts the distinction survives the whole way:
the machine is listed as failed rather than refused, the failing resource
is named in the host's own words, and the machine that did as it was told
is not implicated.
2026-08-30 18:15:14 +02:00
jschoubben e0127df7ce A machine in the mesh builds a module, and the catalogue records it
The chain this closes: a repository exists, the mesh asks for it, a build
machine takes the work, publishes what it made, and the catalogue then
says what the module is, which commit it came from, and — after the
source moves — that it is behind.

Three assertions, against a real broker and registry, because what is
under test is four processes agreeing over a wire:

- the mesh asks, a machine builds, and the artifact is really in the
  registry at the digest the manifest names
- a build that cannot succeed says why and records nothing. A failure
  that is silent is indistinguishable from a builder that is not running
- the source moving makes the catalogue say "behind", and rebuilding
  catches it up

git is now in the base image, with the same reasoning as docker and
wireguard-tools: a machine that builds modules clones them, and a sealed
scenario cannot install anything. Read back from `git --version` rather
than from the package manager — an installed package is not a
capability, and a build machine whose clone fails does so three minutes
into a scenario with the failure reported as a build problem rather than
a lab one.
2026-08-30 04:06:23 +02:00
jschoubben 99b3444b19 Two machines, one mesh, a credential neither end had to be told twice
Everything before this proved a part. This proves the parts meet, which
the project keeps saying cannot be checked any other way.

A bare machine applies the substrate bundle and becomes a mesh — store,
schemas, broker with a certificate it generated itself, control plane
serving. Both machines then join it with nothing but a token. A database
is declared on one and an application on the other, and after a push:

- both ends hold the SAME password, or nothing could authenticate
- it is mode 0600 on the machine that uses it
- it appears in neither machine's stored declaration, neither machine's
  reported state, nor the control plane's database — so it was not
  readable by the broker that carried it or the mesh that sent it
- the consumer is also told where its database is, by a name the mesh
  wrote into that machine's hosts file

The bundle's image references are rewritten to the ones this scenario's
registry serves. A digest belongs to whatever registry serves it, so a
committed bundle names a registry that is not this one — rewriting is
what makes it applicable rather than a placeholder to tidy away.

Two faults found getting here, both fixed in mesh-host: `apply` could not
read a file the bundle could, and the token did not say what the mesh
calls the machine.
2026-08-30 03:05:07 +02:00
jschoubben e89379fef8 Prove a mesh credential becomes a login, against a real database
The mesh generates a password, seals it to the machine that must accept
it, and discards the plaintext — so it cannot tell PostgreSQL to start
accepting it. Something on that machine reads what the host wrote and
makes it true. Everything up to that step is proven elsewhere; this is
where a password either becomes a login or does not.

A scenario with one machine and a database, and six assertions: the
password works, running again reaches the same state and says nothing,
rotation makes the new one work and the old one stop, a departed consumer
loses its login, a role nobody here made is left alone, and a manifest
naming a credential that was never written is refused rather than
creating a login with no password.

Each was confirmed to fail — and only it to fail — with the behaviour
removed from the provisioner: only-creates breaks rotation, no-revoke
breaks revocation, revoking everything breaks the bystander role, and
ignoring a missing credential breaks the refusal.

Two faults in the test itself, both worth recording:

- it checked logins from inside the database's own container over
  127.0.0.1, which PostgreSQL's default pg_hba trusts. No password was
  ever verified. Demonstrated directly: over loopback a deliberately
  wrong password still returns a row. Only the rotation assertion
  noticed, because it is the one that requires a password to STOP
  working — which is an argument for writing that assertion every time.
- the fix then read .NetworkSettings.IPAddress, which docker 29 no
  longer populates. It templates to empty, psql falls back to a unix
  socket that is not there, and every login looks impossible rather than
  misconfigured.
2026-08-30 01:29:46 +02:00
jschoubben 4097ff92c1 Repoint ADR references after HQ consolidated 65 records to 23
Comments naming records that no longer exist now point at the consolidated
record holding their reasoning -- the four lab records are 0016, a test defends
a decision is 0017.
2026-08-28 23:33:46 +02:00
jschoubben 16c13807a9 place: the host — the lab acquires a consumer
The lab raised an underlay and put nothing on it: correct, and useless, because
the thing it exists to test did not exist. Tier 0 now does, so `place: [host]`
works and a raised scenario finally contains something.

The refusal narrows rather than disappearing. A scenario placing a host and a
substrate is told which half is missing, by name — not that `place:` is
unsupported when half of it now works.

Placement reads back rather than assuming. A file arriving is not a host
working, so the binary is run before it is trusted to answer questions, and what
it reports is read from the machine (ADR 0035). The binary comes from an
explicit path, because the declaration design leaves where artifacts come from
open and a search would harden into the answer by accident.

The integration test that matters is the one asserting the host reports the
MACHINE and not the workstation that placed it. A raised VM and this workstation
differ in every capability — root versus uid 1000, a clean init versus a
degraded one, no docker versus docker, no wireguard versus wg0 — so a host
reporting the wrong machine is obvious here and invisible anywhere else.

And the placed host independently confirms ADR 0031: overlay absent on a freshly
raised machine. The underlay suite already asserted that by looking for
wireguard interfaces; this is a second witness rather than the same check twice.

Two tests failed the moment placement worked, which is what they were for. They
defended "there is nothing to place yet" while that was true; the decision
changed, so they change with it rather than being deleted.

Gate: 75 unit, 20 integration.
2026-08-26 00:41:25 +02:00
jschoubben b015068921 Step 2: raise a second scenario, and share the invariants
The suite next door raises one scenario and asks deep questions of it. This one
asks shallow questions of every scenario — the half that was missing, since both
faults found by hand lived in scenarios nothing ever built.

Adds bootstrap-single, the cheapest, and the loop that lets the list grow. Also
adds the second universal invariant: every address a scenario declared is one
the machine actually holds. A machine that came up bare looks identical to one
that came up correctly until something asks it.

Verified to bite rather than assumed: against a live instance, the real
declaration passes and a declaration claiming an address nothing holds fails
with 'anchor declared 192.0.2.99 on hosting but holds 192.0.2.10'.

Integration now runs with --test-concurrency=1. Two files raise real instances,
node --test runs files in parallel by default, and two concurrent runs of this
suite already produced a whole-suite failure once — every test red, from
resource contention rather than from any fault in the code.

Gate: 45.7s -> 60.2s.
2026-08-25 00:29:59 +02:00
jschoubben 715f367147 Step 1: an invariant that holds of any raised scenario
The address collision was found by eye. This is the mechanical form of it: no
two machines hold one address on one segment.

Pure over already-collected facts, so the logic is tested without a hypervisor
— including the cases that would make it useless if got wrong: the same address
on DIFFERENT segments is normal and must not be reported, and one machine
holding an address twice is not two machines.

Asserted against whatever the integration suite has standing, read from the
hypervisor rather than from the declaration. The declaration is what was
accepted, and it was accepted.
2026-08-25 00:21:09 +02:00
jschoubben 2243618f01 Draw a scenario, from the declaration and from the hypervisor
`mesh-lab diagram` renders a scenario as draw.io, from either source, through
one layout — so a difference between what was asked for and what exists is a
difference you can see.

The shape says what a resource is and is fixed per kind. The badges say what is
true about that particular one and come entirely from metadata: translation,
forwardability, mapping expiry, refuses-inbound, container-or-VM, running. The
interesting properties of a network are exactly the ones with no visual
consequence — a translated address looks identical to an untranslated one.

For the live picture to be a record rather than a restatement, raise now writes
down what it applied: a segment's kind, ranges and MTU on the link; a gateway's
translation, forwardability and expiry on the gateway; inbound: deny on the
machine. Every behavioural tag is written AFTER the thing works, never at
creation — a failed raise leaves wreckage standing on purpose, and a picture of
that wreckage must not badge translation the router never got.

The pairing earned itself immediately: drawn side by side, every virtual machine
held no addresses. A container's interface carries the device's name and a VM
names its own, so joining them by name silently dropped one whole class of
machine. Fixed by joining on MAC.

Also brings tests under the typecheck gate, which caught integration timeouts
being passed as a 4th argument and therefore ignored entirely.
2026-08-24 22:53:00 +02:00
jschoubben ca2bbab836 Integration tests, each named for the decision it defends
Reviewed and the criticism was right: 1,072 of 2,128 lines untested, all of
it the half that touches the hypervisor, and no gate. The verification I had
done was real — pings across NAT, TTL counts, ruleset comparisons — and none
of it survived the terminal it ran in, which is 04-ISSUES/005 in miniature.

Ten integration tests against a real hypervisor, each named for what it
defends. ADR 0031: a raised machine carries no overlay, no wireguard, no
mesh config — a scenario that pre-built peering would certify its own work.
ADR 0032: exec is the only way in. ADR 0033: routers are containers while
machines are virtual machines. And the design's claims: raise waits for
usable, snapshots are whole-scenario, NAT hides a private address,
published reaches the machine at the gateway's address.

Mocking the hypervisor is forbidden, so they skip with a reason on a
machine that cannot raise scenarios rather than passing green having
checked nothing.

The suite earned itself on its first run. It found that a snapshot of a
running machine could miss a file written seconds earlier — not stale,
absent — because the write was still in the guest's page cache. That is
exactly the question the lifecycle design listed as open: does a scenario
snapshot need the machines stopped? It does not, but it does need them
flushed. snapshot now syncs every machine before capturing, and the design
records the answer.

The fix buys write-durability, not application-consistency: anything
mid-transaction is still captured mid-transaction, and that is now stated
rather than assumed.

npm run check is the gate — typecheck, 40 unit tests, 10 integration tests.
2026-08-24 22:26:34 +02:00