Commit Graph
253 Commits
Author SHA1 Message Date
jschoubben 33b991b6e0 lab: provider runtime images carry their CLI; prove the private-network shape
build-module-runtime.sh adds psql to the postgres image and mc to the minio image
(their clients shell out to those). provider-on-backend-network asserts redis's
runtime, on the backend's private network, binds the broker via NAT and provisions
a consumer with the mesh's credential — the shape the committed provider manifests use.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 00:52:17 +02:00
jschoubben aafa11756a e2e: a provider creates the resource with the mesh's credential (ADR 0053)
Assigns redis as a provider, puts the contributions and unsealed password the mesh
would deliver in its receives path, and authenticates as the consumer with the mesh's
password — PONG proves the login was created with exactly that password (a
self-generated one answers WRONGPASS), with MESH_SEAL_KEY set nowhere.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 00:27:46 +02:00
jschoubben 6067ec1724 e2e: a running runtime picks up a settings change (issue 009)
Assigns grafana configured by settings, changes the token, pushes again, and
asserts the container was replaced (new id) and the rendered config carries the new
value. Builds the host from source, since the behaviour under test is the host's.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 23:40:13 +02:00
jschoubben 0a414d576a e2e: assigned sonarr + grafana prove the two runtime config paths (ADR 0051/0052)
assigned-sonarr proves the Servarr detection path: the runtime discovers its API
key from the app's config.xml and serves its tools. assigned-grafana proves the
settings path: the operator states URL and token as settings, the mesh merges them
into the module's config file, and the runtime serves from that with nothing in the
manifest. Both green.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 23:08:54 +02:00
jschoubben f093769354 e2e: assigned plex + redis prove the module runtime (ADR 0052)
build-module-runtime.sh generalises the audit-logger runtime image to any module
(mesh-tools + sdk + the module's dist, entrypoints for tools/events/provisioner).
Two scenarios and two tests: assigned-plex proves a tools+events module serves its
tools over a mesh-issued scoped account; assigned-redis proves a provider's runtime
serves tools AND runs its provisioner in the same broker-bound process, provisioning
a grant and emitting its lifecycle event. Both green.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 22:29:53 +02:00
jschoubben 37b16cd0e9 events: the assigned audit-logger, proven in the lab (ADR 0048)
test/integration/assigned-audit.test.ts raises a node into a mesh, assigns it
the audit-logger through the control plane, and asserts the mesh delivered a
scoped amqps account (not the broker's own), the host ran the container, and an
emitted event reached the trail — the delivered credential authenticating is
the proof. scenarios/audit-node.yml is the lean single-node bed that stocks the
runtime image. Passes 1/1 against the real lab.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 02:13:01 +02:00
jschoubben ba7f7998f7 events: an e2e test — an emitted event reaches the audit trail over the mesh's broker
Adds test/integration/events.test.ts: raises first-node (which raises the
broker as tier-1 substrate), runs the runtime+audit-logger against that
broker, emits a module and a node event, and asserts they reach the trail
with their metadata read back from ADR 0047 headers (a pure body), plus that
the durable per-consumer queue and mesh.events.dead exchange exist on the
raised broker — asked of the broker, not assumed.

scripts/build-runtime-image.sh builds the self-contained runtime image
(mesh-tools + vendored sdk + audit-logger) it runs, saved to a tar for
MESH_LAB_RUNTIME. The events path itself is verified; the incus raise is the
part a lab run exercises.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 00:56:39 +02:00
jschoubben 01ca02ad04 A machine works through its queue, and the wait allows for it
With caught-up finally an equality, the last red test turned out to be
telling the truth about something real: declarations queue, the machine
applies them one at a time at half a minute each, and by the twenty-
fifth test it is minutes behind the latest push. 240 seconds was not a
generous bound on one apply — it was an accidental bound on the whole
backlog.

Doubled rather than tuned, and the real remedy filed instead: a machine
asked to be five successive things should become the last one, which
is a decision about the link rather than about this timeout
(novox/hq 04-ISSUES/031).
2026-09-02 01:05:51 +02:00
jschoubben 81e5f31910 settled() asks which, not when
The timestamp comparison lost the race between one test's closing push
and the next test's opening one: the old apply's report landed newer
than the new send and settled() passed for a declaration the machine
had not read. The report names its declaration now, the mesh says
whether it is the current one, and this reads the answer instead of
inferring it.
2026-09-02 00:03:00 +02:00
jschoubben a65115fc81 The forge poll covers both halves, and the cache failure names the container
The provisioner makes the role and then the database; a poll that
waited for the first and checked the second once was racing the gap
between two statements, and lost it once, eighteen seconds into a run.

The cache test's provisioner could not resolve the store's name, which
usually means the store's container never registered it — and the
diagnostics showed only the provisioner's side of that conversation. A
grant that never arrives now prints the container states, the store's
log and the provisioner's, so the next failure names the half that
actually fell over.
2026-09-01 23:20:38 +02:00
jschoubben 5e22e6bc49 The resolve-together test carries the whole catalogue
Eleven modules planned as one set on one machine, which is what a real
node looks like. Two are left out by name rather than silently: the
mesh under test already runs a module called registry and an adopted
workload called umami, and adding the catalogue's manifests would
replace the records of things that are live and assigned — the adopted
umami would suddenly require a database it never asked for.
2026-09-01 23:09:49 +02:00
jschoubben 384e0f4e7e Current is not caught up
settled() returned the moment a declaration was current, because the
sent digest is recorded at send — so both new tests asserted on a
machine still applying, and found containers not created yet and
bindings not written. The eternal-waiting fault had been standing in
front of this gap the whole time; fixing it is what let the tests get
far enough to fall in.

Caught up now means the machine's own last report is newer than what
was sent to it — two timestamps the mesh recorded itself, read from the
`reported` section status now carries. The residual latency between
"reported" and "every container answers" stays with the tests' own
polls, where it always was.
2026-09-01 23:03:05 +02:00
jschoubben 2f2f1d931e The cache edge, proven end to end
A consumer contributes a key prefix and gets an ACL user; the test is
that the grant means exactly what the manifest said, in both
directions: its own keys usable, anyone else's refused by the store
itself, and the flush a tenant must never have refused with them.

Waited for through the store rather than through logs: the user list,
asked with the password the host wrote into the server's own conf file
on the machine — nothing invented, both ends reading what the mesh
delivered.

And the forge is asked on the port the mesh assigned, not the one the
module declared. The old curl aimed at 3000, which was right until
ADR 0038 moved the machine side — a latent break that would have fired
on the first run to get past the settling that used to fail first.

The scenario stocks redis and its provisioner, and the rebuild builds
the provisioner image with the others.
2026-09-01 22:45:05 +02:00
jschoubben 3c9b5a848b The receipt names what was built, not what git says at the end
whatWasTested read the repositories when the run ended, so a commit
landing during the twenty minutes a suite takes was recorded as tested
without ever being in the binaries. It happened: one receipt named a
commit made mid-run, and the verdict it carried belonged to an older
tree.

The heads are read once, right after the build, and carried to the
receipt. A verdict is only worth something attributed to one exact
state, which is the receipt's whole reason to exist.
2026-09-01 22:03:44 +02:00
jschoubben 0d286b7cc8 What review found in the lab, fixed
A segment named "uplink" is refused. The lab claims that name for the
NAT bridge behind `egress: true`, and a scenario wearing it first would
have its egress machines silently attached to an isolated bridge — a
declared key doing nothing, which is the fault this repo exists to
refuse, in the repo that refuses it.

settled() parses inside the try. A truncated status from a struggling
machine was the one shape of bad answer that still threw out of the
wait, and the likeliest moment for one is exactly the machine the poll
is watching. Malformed now counts as "could not ask", like the exec
that times out.

And a sentence on the uplink's UseDNS saying its inertness is
load-bearing: it matters only where systemd-resolved runs, and on a
machine whose modules own resolv.conf the uplink must not outvote the
resolver a scenario is testing.
2026-09-01 21:54:59 +02:00
jschoubben 4a343a2652 A machine is as big as the scenario says, and may reach the world
Three changes, found by one failing test.

The forge failed three runs in a row as "status hangs", and it was
diagnosed twice as contention — real defects, fixed, and not the cause.
The heartbeats told the truth in the end: every exec on anchor crawled
from 15s to 105s, because eleven containers plus a database pull were
running in a 1GiB machine. Starvation presents as whatever you were
doing when the page-outs start, which is why it wore two other bugs'
clothes first.

So machine size is now the scenario's to declare — memory and cpus per
machine, default unchanged. The anchor that carries the whole substrate
is bigger than the laptop that joins it, and the comment on the
scenario says why in terms of what lands there.

`egress: true` gives a machine one extra interface on a lab-supplied
NAT network, addressed by DHCP because the one address a scenario has
no business choosing is on the host's side of the fence. Declared
per machine and off by default: a closed scenario stays the rule
(novox/hq ADR 0016), and the exception exists because a first node
fetches its images before any mesh can serve them — which is now the
tested path (04-ISSUES/029), and a lab that can never reach upstream
cannot prove the bootstrap it exists to prove. The uplink route is
metric-4096, so it never shadows a route the scenario declared. A
detached machine declaring egress is refused, not ignored.

And settled() treats a poll that threw as a poll that missed. An exec
timeout at minute four of a wait is "could not ask", not a verdict on
the machine.
2026-09-01 21:42:18 +02:00
jschoubben f85dbb0713 The registry is added, not built
A mesh that has just bootstrapped cannot build the module that gives it
an artifact store: building publishes to the store, and the builder will
not start without one (novox/hq 04-ISSUES/029).

This test built it and passed, because the scenario's registry was
already standing to receive the push — which is precisely why a real
first mesh would have hit this and the lab never did. A stand-in for
Docker Hub was quietly also standing in for the thing under test.

So the module now names its image by digest, the way the bundle names
the three a first node starts from, and is added as a manifest rather
than built. That is the only path open to a real first mesh, so it is
the path this walks.

Its skip on MESH_LAB_BUILDER goes with it. Nothing in the test needs a
builder any more, and a skip that names a thing the test does not use
sends the next person to look in the wrong place.
2026-09-01 21:17:59 +02:00
jschoubben 2b3a30619b Stop raising a second scenario to test the first three tests
The canary walked one path on one machine — a mesh comes up, a module
lands, a consumer gets a credential — and stopped the run if it broke.
That path is exactly what the first three tests of the long run walk,
and the long run finishes them about 160 seconds in.

So the gate cost a whole scenario on every passing run to save roughly
45 seconds on a failing one. A scenario is three machines, one of them a
registry that boots a kernel in order to serve files, which is where the
two minutes went.

The test file stays and still runs when it is named. What is gone is
raising it on the way to everything else.

Measured rather than argued: the canary's scenario took 116s of which
60s was standing up a registry, and the run reached the same assertions
without it.
2026-09-01 21:07:07 +02:00
jschoubben 144362be37 A poll that could not ask has not been answered
`settled` used `must`, so a failed exec ended the wait as though the
machine had reported a failure. It had reported nothing: the control
plane is a container on the node being polled, and while that node
applies a declaration an exec into it can lose its stdout fifo to
containerd. The run then blamed the mesh for a question that missed.

Could not ask and asked, and the answer was bad are different facts, and
only the second is the machine's. A failed poll now keeps the reason and
tries again; the timeout reports whichever came last, so a control plane
that is genuinely unreachable still fails the test — with the reason
rather than with a stack trace.

Every five seconds rather than every two. Each poll is an exec into a
container on a machine that is busy applying, and thirty times a minute
was competing with the apply rather than observing it.
2026-09-01 20:25:54 +02:00
jschoubben 2bf5846bf2 Wait for the machine before asking what it is running
The forge test read `status` the instant `push` returned and concluded
the machine was fine. It was describing the apply before this one.

`push` sends and returns — it prints "sent N resource(s)" and the
machine applies afterwards. So every assertion made immediately after
one is racing it, and this race lost quietly: no failure reported, and
a container that did not exist yet read as a container that would never
exist.

`settled` asks the mesh, in its own terms: a node is caught up when it
is neither waiting for what it was sent nor wrong about what it applied
— the two questions `status` already answers, read as JSON so a test is
not parsing a report written for a person. A machine reporting a failure
ends the wait immediately rather than at the timeout, because it will
not become right by being waited for.

The container assertion now also prints the plan. A container missing
because the mesh never asked for it and one missing because the machine
could not make it are one sentence and two entirely different faults,
and the plan is what separates them.
2026-09-01 19:34:23 +02:00
jschoubben e99bffc93e Ask the machine what it did, before asking what it produced
The forge test pushed and then waited for a database login. When the
containers were never created at all, it reported "no login was created"
— true, and silent about why. Two hundred and thirty seconds spent
proving something downstream of the actual failure.

A push being accepted and an apply having worked are different facts,
and this test depends on the second. It now reads what the machine says
about itself, and whether a container exists, before it starts waiting —
and carries the host's own log into the failure either way.

The suite otherwise passed 24 of 25 on this run, which is the first time
the forge reached a clean attempt with nothing upstream blocking it.
2026-09-01 16:44:21 +02:00
jschoubben eba436b6b7 Reach one scenario from the workstation, by name
A scenario is a closed address space: two raised from the same
declaration hold the same addresses and never meet, which is what lets
two run at once and why the lab talks to machines through the
hypervisor rather than over IP. Reaching in from outside breaks that, so
it is opt-in, one scenario at a time, and reversible.

`connect` takes an address on the scenario's public link and writes a
resolver rule answering everything under each machine's name.
`disconnect` gives both back. `connected` says what is true right now,
for somebody who cannot remember.

It refuses rather than guessing when more than one scenario is standing
— the failure being avoided is not an error but one scenario's traffic
arriving in another. It also refuses when a machine's name is already
answered here for something real, because connecting would point that
name at the lab, and the damage would land on the real thing.

Names answer with the segment address rather than the overlay one.
Inside the mesh a name gives a machine's private address; from here that
would need this workstation on the overlay, which is a much larger door.
The segment address reaches the same machine and the same ports, which
is what opening a board in a browser actually needs.

Proven against a live two-node scenario: registry.internal:5000/v2/
answered 200 from this workstation, and so did a wildcard name under the
same machine. Disconnect put the address back, stopped answering, and
left the real mesh's own names alone.

One thing measured rather than assumed: it restarts dnsmasq instead of
reloading it. A reload is SIGHUP, which re-reads the hosts file and
clears the cache but not the configuration — the rule was written, the
reload reported success, and nothing resolved. The daemon's start time
was nine days old afterwards.
2026-09-01 16:42:26 +02:00
jschoubben 6591a2e040 A canary first: one machine, one path, three minutes
Suggested by Jochen, and it paid for itself on its first run.

A suite that takes forty minutes is a suite you hear from once a day.
Every fault found today would have shown up in the first three minutes
of it — a module pinned to an image that does not exist, a consumer
given a password and no name to present with it, a credential file
nothing could read, a search for a password that read the password as an
option. The other thirty-seven minutes proved things that were already
working.

So this runs first, on one machine, with the three images the mesh needs
for itself. It walks one path: a mesh comes up, a module lands, and a
consumer gets a credential it can actually use — the name to present,
the address, the port, and a password only the host could put there.
Deliberately not a smaller copy of the full suite: that path is where
everything went wrong, and a canary checking many things shallowly is a
canary whose failure nobody can read.

`suite` runs it and stops if it dies, saying why rather than leaving
somebody to wonder what the missing thirty-seven minutes would have
said. Skipped when the caller named its own files.

It measured 164 seconds against forty-odd minutes, and failed three
times on its first run for one reason: applying the bundle raises a
control plane but does not tell it a machine exists. I had left out
enrolment, and the long suite would have taken forty minutes to say so.
2026-09-01 16:24:19 +02:00
jschoubben 44d3e53bcb Three failures, all mine, all worth having
**A password beginning with a dash broke the search for it.** The
credential test greps the machine's own files for the delivered
password; this run's password started `-S`, so grep read it as an option
and refused the whole invocation. The test compared the usage message
against "0" and reported the password as leaked. That is the worst way
for a search to fail — it says it found something. Fixed with `-e` and
`--`, which is what those exist for.

**The planning test could not redirect what the scenario does not
serve.** Rewriting an image reference only works for repositories the
scenario's registry actually holds, and the object store's provisioner
was not stocked — so that module kept its placeholder and the refusal
fired, correctly. It is stocked now, so all five are planned again. The
skip path stays for anything genuinely unserved, and says which module
and why: a planning test quietly covering four instead of five is the
false coverage this suite exists to prevent.

**The forge failed because of the one above it.** The planning test
threw before its cleanup could run, leaving a module assigned that
refused the next push, so no database container was ever created.
Yesterday's fix moved that cleanup where a failure cannot skip it — but
`after` still only unassigns what was assigned, so it now tracks what
actually got added rather than what was intended.
2026-09-01 16:16:57 +02:00
jschoubben 2c7e69f4db Plan what could actually run, and put the machine back either way
Two faults, both found by the guard that now refuses a placeholder
digest on its way to a machine.

The planning test added the manifests exactly as they sit on disk, which
includes an image the mesh builds — and that image has no digest until
it is built, so the file legitimately carries a placeholder. The test
was therefore planning something that could never run, which is the
whole complaint. It now points the references at this scenario's
registry first, exactly as the forge test does.

The second is worse and more ordinary. Its cleanup was the last
statement in the test body, so the first failure skipped it and left
five modules assigned. The next test's push was then refused by a module
this one had abandoned — a failure that reads as a fault in the test
that was working. Cleanup that only runs on success is not cleanup, so
it moved to `after`, where a failure cannot skip it.
2026-09-01 15:54:56 +02:00
jschoubben 55022edc1e Run the forge, on a database the mesh gave it
Everything up to now stopped at composing a declaration. That proves the
control plane and the host agree, and proves nothing about whether the
thing described works — which is how five modules sat pinned to images
that did not exist while parsing and resolving perfectly.

The forge is the right one to run first. It needs a database from
another module, a password it did not choose, and a connection string it
could not have written itself: the address and port come from what the
database serves, the user name from what the mesh decided both ends
would call it. If any of that is wrong it cannot start, and nothing else
in this suite would notice.

The test checks the chain in the order it has to happen — the login
exists, the database it owns exists, the forge answers, and its log does
not say authentication failed. That last one matters: a forge that
started and could not reach its database would still answer on its port.

Rewriting image references is now shared rather than copied from the
bundle, which had the same problem first. A digest is not knowable until
something is built, and when it is, it belongs to whichever registry
served it — so the text says which image and the scenario says which
copy. Matching is on the repository, with a test that a repository
ending in another one is not half-replaced.

Also makes the planning test put the machine back. Tests here share one
mesh, so the five modules it assigned were inherited by whatever ran
next; harmless while nothing pushed, and not harmless now.
2026-09-01 15:39:30 +02:00
jschoubben bb14ecb7e0 Say what the lab is doing, while it is doing it
novox/hq 04-ISSUES/024. A run stalled for thirty-five minutes and said
nothing. The cause was a link systemd was still configuring, three
layers down inside a `docker load` blocked on a socket — and every one
of those layers knew what it was waiting for. None of them said so.

Three decisions, each doing work.

**Every external command is logged, at the three places that run one.**
Ninety-seven call sites reach a hypervisor or a container runtime
through three wrappers, so instrumenting the wrappers covers all of them
and nothing has to remember to log.

**A command still running says so while it runs.** A line before and a
line after tells you nothing until the after arrives, which is exactly
the case that matters. Anything outstanding past fifteen seconds reports
itself with how long it has been going. It is reported as still running,
not as stuck — which it is is not knowable from there, and a log that
calls a slow step a hang teaches people to ignore it.

**It goes to a file, written synchronously.** Node block-buffers stdout
when redirected and a test runner buffers it again, so a console log can
sit minutes behind. `appendFileSync` cannot lag.

Two things this found in itself while being written, both the same shape
as what it exists to catch:

A question that answers no is not a fault. Half the lab's commands are
questions — does this network exist, is the agent up yet — and they fail
constantly while a scenario comes up. Logging those as faults filled a
healthy run with ✗, which is how you end up ignoring ✗ when one is real.
They are recorded quietly now, and still recorded.

And `around` skipped its own wrapper when a step's level was below the
configured one — taking the failure line and the heartbeat with it. The
two things worth having at a low level were the two that vanished at
exactly the level somebody would use. The gate belongs in `write`.

Also unsilences the four call sites that passed a callback throwing
everything away, including the one the stall sat in, and tees `raise`'s
progress into the file whether or not a caller asked to see it — the
end-to-end test passed no callback, so the one run that mattered
reported not a single step.
2026-09-01 10:43:38 +02:00
jschoubben 5137720aa7 A path is not a secret, and the check said it was
The last failure of the run: `MESH_BROKER_FILE=/var/lib/mesh/builder/broker`
reported as "something secret-shaped, which the broker would see".

`/` is in the base64 alphabet, so any absolute path of 24 characters or
more matched the pattern meant to catch a sealed value. An absolute path
is a *reference* to a secret and naming one is the whole design — the
mesh delivers a credential as a file and a module says where.

Excluded explicitly rather than by loosening the pattern, and checked
both ways: a real sealed value and a base64 blob are still flagged, a
relative path still is, only an absolute path is passed over.

Worth the words in the comment. A check that fires on the right shape
for the wrong reason is worse than none — it is the one that gets
suppressed, and then it is not there when it is right.

It surfaced now because tests in this file share one mesh: the builder
was assigned by an earlier test and appears in this one's declaration.
2026-09-01 09:52:28 +02:00
jschoubben df4dd406a9 A redirected log lags; do not diagnose a stall from it
Node block-buffers stdout to a file, so a run log can sit unchanged for
minutes while the run is fine. Read that way twice today — the second
time straight after fixing a real stall, which is the worst version of
it, because a buffering artifact then reads as the fix having failed.

The machines are the source of truth and answer immediately. Written
down with the two commands that settle it.
2026-09-01 09:40:45 +02:00
jschoubben 3503ad990b The registry is addressed the way every other machine is
novox/hq 04-ISSUES/024. The registry machine had its address set with
`ip addr add`; every other machine gets a systemd-networkd unit. That
one difference stalled the lab indefinitely.

An address set by hand leaves networkd waiting to configure a link it
was never told about, so the link sits at `configuring` for ever.
`systemd-networkd-wait-online` has TimeoutStartUSec=infinity, so
`network-online.target` is never reached — and Docker is ordered after
it. `docker load` then blocked on a socket whose daemon was queued
behind a target that would never come.

Measured before and after on the same scenario: stuck with five pending
systemd jobs and `docker` inactive; now `enp5s0 configured`, `docker`
active, no jobs, and the whole raise completes in 87.5s.

The guess in the issue was wrong, and it was wrong in the usual way —
stocking had just been changed, so stocking looked guilty. Stocking
takes 34s and always did.

Two things that made this cost hours rather than minutes are fixed with
it. Placing an image now waits for the container runtime to answer and
refuses after 120s naming what systemd is waiting on, so a stall becomes
a failure that says why instead of three stacked timeouts totalling 35
minutes. And the end-to-end test passes `onProgress`, so a raise says
what step it is on — it printed nothing at all until it finished, which
is why 35 minutes of nothing read as a slow test.
2026-09-01 09:37:16 +02:00
jschoubben c30d71e249 Point the builder at build/, which is ignored
Pointed at the repository root, the binary it writes is 12 MB of tracked
artifact. Said in the example rather than left to be discovered by a
On branch initialization
Your branch is ahead of 'origin/initialization' by 43 commits.
  (use "git push" to publish your local commits)

Changes to be committed:
  (use "git restore --staged <file>..." to unstage)
	modified:   README.md that looks wrong.
2026-09-01 03:21:53 +02:00
jschoubben e83e5ab24e Prove a hole in a config file is filled, on a real machine
Nothing did. The sealed-placeholder substitution and the bound-value
substitution were each covered by unit tests in the repository that
performs them, and the two expressions that find the holes live in
different repositories — so both sides could agree with themselves and
disagree with each other, and the first thing to notice would be a
program connecting to a host called "${bound:postgres-database:at}".

So the consumer in the credential test now ships a configuration file
with four holes in it: the address and port from what the provider
serves, the name to present from what the mesh decided, and the password
sealed. The mesh fills the first three before sending, the host opens
the credential and fills the last on the machine, and the test reads the
file off the machine and checks that the password in it is the same one
the credential file holds — and that no ${ survived.

This is the only place those two mechanisms meet a real host.
2026-09-01 03:15:22 +02:00
jschoubben facd100baf Rebuild the object store's provisioner image too
It gained a Makefile target today; a target the rebuild does not run is
the stale image this file was just fixed for.
2026-09-01 03:12:55 +02:00
jschoubben eb02b8fe6b Assert the whole of what keycloak is given, not one file's name
The credential moved: the sealed password is a password alone, at
`.secret`, and `database.env` is now the connection keycloak could not
have written — address and port from what the provider serves, user name
from what the mesh decided both ends would call this consumer.

So the test asks for both, and for the seam between them: the password
is still a hole, the sealed value travels beside the file that needs it,
and no ${bound:...} survives as a value. That last one matters most —
a placeholder written through would be read as a hostname, and the
failure would name neither the module nor the mesh.
2026-09-01 03:06:59 +02:00
jschoubben 51af4307a9 Rebuild every image the lab runs, not only the control plane's
A run today had a control-plane image built that minute and a
provisioner image built the day before. The rotation test failed against
a real database and it looked exactly like the change under test being
wrong — the provisioner was creating logins by a naming rule that had
been replaced hours earlier.

It was the rebuild. It covered `make image` and the builder binary and
none of the three other image targets, all of which the suite runs.

This is the same fault the builder line was added for, one target along,
and the comment there already names the precedent: building one and not
the other is the eleven-hour-old binary. A rebuild that covers most of
what a run uses is worse than one that covers none, because the run that
follows it is believed.

The test names each target rather than counting them, because what goes
wrong is a target that exists and is not run, and a count would not
notice.
2026-09-01 03:04:51 +02:00
jschoubben cb0edcee8d The end-to-end test reads the grant where the mesh now writes it
Two assertions in the full-mesh test encoded the old naming: the grant
file read back from the provider, and the PostgreSQL role the real
application logs in as. Both are named after the consumer now, and a
consumer is a module on a machine.

These are the two that matter most in this file — it is the only place
where a real application authenticates against a real database with a
password the mesh delivered and cannot read, so they are what would have
caught the naming going wrong end to end.
2026-09-01 02:48:24 +02:00
jschoubben e7f4a49e40 A grant file and a provisioned login name the module too
novox/hq 04-ISSUES/022: a consumer is a module on a machine, not a
machine. The mesh now writes <node>.<module>.secret and the provisioners
name the role and the access key after both.

The fixtures here write what the mesh writes, so they move with it —
that is the whole point of them, and a fixture that kept the old shape
would agree with the bug rather than catch it.

The object-store assertions are the ones that mattered most: one store
holds every bucket behind one endpoint, so isolation is a policy rather
than a property. One access key per machine meant every module on a node
shared it, and the policy confining each consumer to its own bucket
confined none of them.
2026-09-01 02:45:40 +02:00
jschoubben 0ffb24ff5d Write down what a run has to be pointed at
Reconstructed from the source twice now, which is 04-ISSUES/005 in its
own README: a test whose artifact was not pointed at skips rather than
fails, so an unset variable is a green run that proved nothing. The
first attempt today reported "skipped 24" and left a receipt claiming
zero of everything — working exactly as designed, and indistinguishable
at a glance from a suite that had nothing to do.

Also records the two things that cost time either side of it: `check`
says which variables are missing before a long run rather than skipping
quietly, and a heredoc into `newgrp` runs the suite as a child of a
shell that immediately exits, so it needs `setsid nohup … &` or it dies
with the shell that launched it.
2026-09-01 02:44:32 +02:00
jschoubben a516ee847b Adopt a third-party workload, and keep a mesh between runs
**The adoption.** Software nobody here wrote, taking its credentials the
way such software does — from its environment — and needing two
containers that reach each other by name. The first module that could
not have been declared this morning: it needs the network shape and it
needs a sealed value to reach a container's environment.

Its password is accepted rather than generated, which is the whole shape
of an adoption: a service that already exists keeps the credential it
already has. Asserted properly — a wrong password is refused by the same
database, so the passing case means something.

**The warm scenario.** A mesh kept between runs and returned to, which
turned twelve minutes of bootstrap into thirty seconds of restore. Off
unless asked for: a run that is meant to mean something raises from
nothing.

Its guard fired for real during this work, unprompted — a mesh-host
commit landed and it refused the stale base, naming both commits, rather
than passing tests against yesterday's binary. That is 04-ISSUES/005's
rule one level down.

Three things the guard learned the hard way and now handles: a snapshot
captures disk and not memory, so the host is restarted after a restore
and asserted to have come back; the stocked image digests are worked out
while raising and a restored instance never raises, so they are kept;
and comparing only the repositories this run can see clears the ones it
cannot, so both directions are compared.

The one real bug behind five failed attempts was in mesh-host and it
reported itself precisely: a network shape the language had and no host
implemented. Everything else was scaffolding of mine.
2026-09-01 01:20:14 +02:00
jschoubben bfcbee49e9 A public name, against a real ACME server
The other half of the certificate split: the mesh's own authority
certifies internal names, and a name reachable from outside needs one
the world already trusts. Against a real server rather than a stub,
because what is under test is whether an order, a challenge and a
handshake agree, and a stub would be told to agree.

One assertion passes and one fails, and the failure is filed as
novox/hq 04-ISSUES/020: the authority issues a certificate and the
client never collects it. Kept as a failing test rather than deleted or
skipped — it is the reproduction, and it proves everything up to the
last hop.

The passing one is the guard that matters day to day: no certificate is
ordered for a name the mesh does not route, so a scan cannot spend an
account's rate limit.

The failure output gathers both sides before asserting. The first
version reported only what the proxy said, which made a server-side
question unanswerable — "the client never spoke to it" and "it refused
what the client said" are different faults with nothing in common.
2026-08-31 21:40:16 +02:00
jschoubben 2096d0b2a1 Prove a bucket is provisioned the way a database is
Seven assertions against a real store, the important one being that a
consumer cannot reach another consumer's bucket — isolation here is a
policy somebody wrote rather than a boundary the product has.

The revocation test stages its own precondition. The first version
asserted a key left by an earlier test, and the rotation test had
already revoked it two tests early: the behaviour was correct and the
test was measuring residue. Its precondition assertion is what caught
that, rather than it passing green having verified nothing.
2026-08-31 17:51:12 +02:00
jschoubben 61f864a274 Name ADR 0007 in the connectivity tests that defend it
Certificates, filtering, hub filtering and wildcard resolution were all
proven here without naming the decision they defend. novox/hq ADR 0017.
2026-08-31 17:33:34 +02:00
jschoubben 17e7132532 Name the engine in the provisions these tests declare
Follows novox/hq ADR 0027: a consumer is written against PostgreSQL,
not against a database.
2026-08-31 17:12:46 +02:00
jschoubben 033ad7ec69 A run rebuilds what it tests, and leaves a receipt saying what it covered
The danger is not that the suite breaks. It is that nobody notices it
stopped running (novox/hq 04-ISSUES/005). The harness this replaces had
not built for two and a half months and nothing said so — and this suite
needs a hypervisor, so it inherits exactly that: it runs when somebody
remembers, and remembering is not a mechanism.

So running, recording, and rebuilding are one act:

- the host binary, control-plane image and builder are rebuilt from
  source first. The last two both parse manifests; building one and not
  the other left a binary eleven hours old refusing a field the mesh had
  just renamed, found by a full run.
- a receipt lands in XDG state — outside git, because the question is
  whether *this machine* has run it, and a receipt in git would be a
  claim about everybody's machine made by whoever committed last.
- `last-run` judges it and exits non-zero when it no longer counts.

Three faults found by running the thing rather than reading it, each now
held by a test confirmed to fail without it:

- counted() passed every test while parsing nothing. The runner colours
  its summary even into a pipe; the fixtures were clean text that had
  been imagined rather than captured. A fixture that agrees with the
  mistake proves the mistake.
- a receipt for `suite test/lastrun.test.ts` was indistinguishable from
  one for the real thing — 005's own symptom, rebuilt inside its remedy.
  The receipt now records what ran.
- a tree with uncommitted work reported the bare commit, claiming
  coverage of code nobody can check out. Nothing else could tell: the
  hash is identical either way.

Proven on real machines: 22/22, against all three repositories.
2026-08-31 15:02:19 +02:00
jschoubben e1317c9a69 Gather the evidence however the resolver test fails
Three times a diagnostic has not run because the thing before it threw: `must`
on a command that exits non-zero, and then a query that hung long enough to
take the harness's own timeout with it — which arrives as an error with no
evidence attached rather than as a failed assertion.

A thirty-second test costs fifteen minutes to re-run, so the evidence has to be
gathered whichever way it fails. One helper, used by every assertion here, and
the query is bounded on the machine rather than by the harness: a query that
hangs is a result, not an accident.
2026-08-31 13:54:04 +02:00
jschoubben a83dd3d00a A module says own-secrets, not needs 2026-08-31 13:47:21 +02:00
jschoubben a83f6b7f8e Assign the resolver configuration module that suits the machine
These machines run systemd-resolved, so resolv-conf would fight it over the
file. Both claim the-resolver-configuration so that assigning the wrong one is
refused rather than fought over — and the test was picking the wrong one.
2026-08-31 13:39:58 +02:00
jschoubben a6480bbd88 Say what holds the address, rather than that it is held
dnsmasq cannot bind 127.0.0.54 and the module's own comment claims that address
is free. Which of those is wrong is the question, and naming the holder answers
it — assuming would be issue 012's mistake, where two things changed and the
plausible one was blamed.
2026-08-31 13:24:40 +02:00
jschoubben 89b6dd6080 Print why the resolver did not start, instead of that it did not
`systemctl is-active` exits non-zero for a unit that failed, so `must` threw
before the assertion carrying every diagnostic — and the run said only
"failed". The journal, the config, what the mesh wrote and resolv.conf are all
things the next run should not have to be re-run to see.
2026-08-31 13:11:02 +02:00
jschoubben 20690964f1 Prove a service is reached by a name under the machine it runs on
Through the path an application actually takes — nsswitch, files, then DNS —
because the resolv.conf module is half of what is being tested and only that
path goes through it. Asking a server directly would prove less.

Both machines resolve, from their own copy: a mesh where one machine answers
for all of them stops resolving when that machine does, which is the
arrangement this design refuses everywhere else.

The manifests are read from mesh-control's examples rather than written here,
so what is proven is what ships. And dnsmasq joins the base image, read back
through --version like the others: a machine that cannot answer names applies
the resolver data, reports success, and resolves nothing.
2026-08-31 12:55:59 +02:00