Commit Graph
98 Commits
Author SHA1 Message Date
jschoubben 71bea08f3b anthropic-bed: prove the host unseals the refresh token, no node-key stub
The bed follows the reworked flow: the manager module seals the refresh token to the node's
PUBLIC key, the HOST unseals it and mounts the cleartext at the manager's bound path, and the
refresh reads that cleartext -- no fake node key pair is mounted any more, the host uses its
own real sealing key.

  - the manager is a model-access holder deployed first, so its bound facts (carrying the node
    public key) are delivered; the consumer is added only once an access token exists to seal.
  - adopt reads the node public key from the bound facts; the test asserts the host mounts the
    cleartext refresh token for the manager, and that it reaches nowhere on the consuming node.
  - the refresh_grant assertion reads { sealed, manager_key }.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 01:55:28 +02:00
jschoubben 6be5072565 anthropic-bed: prove model-access refreshes on the manager node
A lab bed for Phase C of model-access (ADR 0050), OAuth endpoint stubbed.
It drives the real runtime images through the whole flow: the manager
seals a refresh token at rest and opens it on the manager node alone,
mesh-control is handed only the access token and an opaque re-sealed
envelope via licence submit-refresh, and the consumer writes an
access-token-only credential. Asserts the refresh token -- original and
rotated -- is nowhere on the consuming node and only ciphertext in the
control plane's database.

build-module-runtime.sh also compiles adopt/refresh/apply/usage
entrypoints. Stubbed and flagged: the vendor endpoint, the manager node's
private key (mounted; a host capability to deliver it does not exist
today), and the submit transport (the test invokes the CLI on the
manager's output).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 01:00:37 +02:00
jschoubben f04c4ba3be lavinmq-bed: GREEN — AMQP provider provisions a require-only consumer's vhost (mint fix)
Two-node: substrate broker on anchor, lavinmq provider + amqp-ping consumer on laptop. The
consumer gets its scoped vhost+user, connects, and round-trips a message. Requires the
mesh-control require-only-mint fix. Diagnostic removed now it's green. SUITE_EXIT=0.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 23:20:59 +02:00
jschoubben 73bd086193 lavinmq-bed: reproduces a credential-mint gap for require-only consumers of a parameterless provision
The lavinmq AMQP provider comes up and serves, but its receives file has given:[] — the
consumer amqp-ping (requires amqp, contributes nothing, as a parameterless provision like
redis-cache takes no per-consumer payload) is never minted a credential, so the provisioner
creates no vhost. Diagnostic in the test dumps the empty grants file + the provisioner log.
This is the resolve/plan mint path (grantsFor -> SecretsFrom), ADR 0048 territory, and it
likely affects redis-cache consumers the same way. Preserved for a focused fix; not merged.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 16:09:01 +02:00
jschoubben 8e494c0e78 Add lavinmq-bed: a two-node amqp provider + consumer bed
A lavinmq provider and an amqp-ping consumer ride laptop while the
substrate's own broker owns 5672 on anchor — the twin of two-node-db.
lavinmq is the mesh's control broker AND a user-facing capability, so a
provider must publish 5672 for its consumers and cannot share a node
with the control broker that already owns it; the split unblocks the
chain single-node.

The bed proves, layered: the run-once bootstrap computed the admin hash
and wrote the broker config before the broker started (ADR 0052); the
service and both runtimes are up and stable; each module got its scoped
broker account on the substrate broker; the provisioner created the
consumer's vhost AND user, both named for the derived login; and the
consumer connected to that vhost with the mesh-minted password and
round-tripped a message. The consumer uses ${bound:amqp:as} for user
and vhost both, and the provider's serves carries the port.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 15:50:33 +02:00
jschoubben 6cd595d403 route-forwarding: GREEN — slug hello-web so it resolves; Host-routing + withdrawal proven
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 15:16:56 +02:00
jschoubben f03647d605 Add route-forwarding lab bed proving the route grant end to end
A scenario and integration test assign route-proxy (provider) and hello-web
(consumer) on one node, then assert a request to the consumer's name -- sent to
the proxy -- is forwarded to the workload and returns its answer, and that
unassigning the consumer withdraws the route so the same request stops working
(the proxy replaces its table rather than merging). Modeled on
mesh-grant-end-to-end and schedule-tick: module add, assign, one push, settled,
with no module issue (route-proxy needs no scoped account).

build-route-proxy-image.sh compiles the Go proxy from
mesh-control/examples/route-proxy into mesh-route-proxy:development for the
scenario to stock. This bed proves route-forwarding over plain HTTP;
public-ACME TLS is proven separately by certificates.test.ts against a real
ACME server.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 15:07:23 +02:00
jschoubben 76ecbaed85 lab: a scheduled container fires on its cadence without gating (ADR 0053)
The bed that proves the scheduled-container primitive end to end. schedtest
is the thinnest carrier of ADR 0053: one container marked
schedule: "* * * * *" that appends a timestamp to a mounted data dir each
time the host fires it -- no service, no listener, no provisioner, no
runtime, no tools, no events.

The three claims it proves, from the ADR's "How each claim is checked":
installing the schedule leaves the node current WITHOUT a run (baseline
captured right after settled, the deliberate inversion of run-once); the
container fires on its cadence (a line beyond the baseline within ~150s);
and it recurs (a second line on the next minute -- cadence, not a one-shot).

schedtest serves and consumes nothing and carries no runtime, so it is not
issued a broker account: module add -> assign -> one push is the whole
sequence, no module issue. The tick image is a bare alpine served by the
scenario's registry by digest.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 14:20:54 +02:00
jschoubben a9ecce25cb Two-node DB-consumer bed + a scenario disk field
The GREEN multi-node regression bed that proves the DB-consumer gate: substrate/control on
one node, postgres+redis providers and baserow+letta consumers on another, each consumer
getting its own credential and its own mesh-named database across the overlay. Requires the
mesh-control provider-seal-key fix and the mesh-catalog db-name fix.

Includes a general lab capability: a machine 'disk' field sizing the VM root disk (a broad
install exhausts the pool default and the host fails mid-apply with 'no space left on
device'). The bed sets 60GiB.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 13:48:26 +02:00
jschoubben e703a94fcd tools-confluence: the tools-only bed for confluence (serves 3 tools, no creds)
Sixth green regression bed. Same shape as tools-gitlab: a runtime-only module comes
up under the mesh, serves its full tool surface with no valid credentials (the Servarr
lesson), stays up, binds its serve queues, and gets its scoped broker account.
SUITE_EXIT=0.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 01:54:21 +02:00
jschoubben c9619a17f4 Add tools-gitlab lab scenario proving the tools-only install
Prove gitlab — the exemplar tools-only, outbound-only external-SaaS
integration — installs: a scenario assigning gitlab to one node, and a test
asserting the mesh-runtime-gitlab container comes up and stays up, logs
[mesh-tools] serving 23 tool(s), binds its serve queues on the broker, and
gets its own scoped account — all with NO valid GitLab token, the case the
Servarr lesson is about.

No gitlab arm is needed in build-module-runtime.sh: gitlab speaks HTTP and
needs no extra CLI in the image.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 01:33:52 +02:00
jschoubben 62e479b9fe catalogue-mqtt: prove the run-once primitive end to end
A bed that assigns mosquitto and asserts the run-once step seeded dynsec
before the broker: the bootstrap ran to completion (not left running), the
seed is on disk owned by the broker's uid, the broker is up and stable
(it crash-loops against an unseeded store, so a stable broker is the proof),
and the node reached current. On top, the seeded admin authenticates over
MQTT and the provisioner grants a scoped client a consumer connects with.

build-module-runtime.sh gains a mosquitto arm (install mosquitto_ctrl from
the mosquitto package — it is not in mosquitto-clients on bookworm, and a
musl binary from eclipse-mosquitto would not load) and compiles the module's
bootstrap/index.ts entrypoint.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 00:33:41 +02:00
jschoubben fc36fb5b51 catalogue-media: two modules accessing one operator-owned directory co-resolve (ADR 0051)
sonarr and radarr both access /services/media/downloads — the exact duplicate path
the resolver refused before novox/hq ADR 0051 (04-ISSUES/036, 012). Each now declares
it as an `access`, not a `directory` resource, so the pair co-resolves and one push
configures both. The operator provides the shared media dirs before apply (the host
refuses an absent access); the bed creates them after enrol and before the push.

Proves: the push is not refused, the node converges once, both modules' server and
runtime containers are up, and both server containers mount the same operator-owned
spool.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 23:02:41 +02:00
jschoubben 571a6cd6fd WIP: catalogue-apps install bed (mongodb, unifi, marrytts)
Adds the mongodb runtime CLI (mongosh) to build-module-runtime.sh, a
catalogue-apps scenario, and its install test. Proven so far: the ADR-0054 slug
applies and the mesh accepts the push (mongodb consumer identity mesh_anchor_mongo
fits). NOT green: the node applies but never reaches 'current' within 1200s — a
persistent reconcile divergence (applied-but-never-current, no crash), likely a
module declaring a resource its container mutates (issue-011 class). Needs live
VM inspection to name the module. Not merged.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 15:30:50 +02:00
jschoubben ab04dd814f Add catalogue-small co-residence bed (four modules, one push)
Raises the first-node substrate and assigns postgres, minio, redis and
plex to one anchor in a single push, proving they resolve and come up
together on one node. postgres and minio each get a consumer that
connects with a real granted credential.

redis follows the corrected provider contract (ADR 0048, issue 032): its
runtime reconciles the contributions the mesh delivers at MESH_RECEIVES
and creates each consumer's ACL user with the mesh-minted password,
sealing nothing — no MESH_SEAL_KEY, no *.grant.json/*.credential path.
The provisioning proof authenticates as the consumer with the mesh's
password (PONG), matching the green provider-uses-mesh-credential bed.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 14:14:56 +02:00
jschoubben cbd9b647ec e2e: unblock the minio grant — the consumer declares a slug (ADR 0054)
Issue 010 fixed: bucketuser declares slug `bkt`, so its identity mesh_anchor_bkt (15)
fits an S3 access key where mesh_anchor_bucketuser (22) did not. Unskips the test.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 02:51:53 +02:00
jschoubben b7ba7534af e2e: the whole grant for an S3 bucket (skipped — blocked on hq issue 010)
Mirrors the postgres bed for minio: a provider (runtime carries mc) + a consumer
requiring s3-bucket, proving the consumer reaches its bucket with the access key and
secret the mesh delivered. It surfaced a real limit: the mesh derives `as` =
mesh_<node>_<module> (22 chars), and an S3 access key is capped at 20, so minio refuses
the service account. The test is correct and skipped pending 04-ISSUES/010, not worked
around.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 01:58:01 +02:00
jschoubben de7f3b9c05 e2e: the whole grant for a database — a consumer connects with what the mesh delivered
Assigns a postgres provider (its runtime carries psql) and a module that requires
postgres-database; the mesh mints one password, postgres's provisioner creates a role
and database under the mesh's login with it, and the consumer connects to its database
with the delivered credential (a password-checked connection) — select 1. Nothing placed
by the test. The postgres half of the per-backend provider proof (ADR 0052/0053).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 01:47:57 +02:00
jschoubben 617b5577f7 e2e: the whole grant, mesh-driven — a consumer authenticates with what the mesh delivered
Assigns a redis provider and a module that requires redis-cache; the mesh mints one
password, seals a copy to each end, writes redis its contributions and the consumer
its bound file, and the host unseals each side. redis's provisioner creates the ACL
user under the mesh's login with the mesh's password, and the consumer's delivered
credential authenticates (PONG). Nothing is placed by the test — the provider/consumer
contract (ADR 0053) working as one thing, no shared key anywhere.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 01:14:48 +02:00
jschoubben 33b991b6e0 lab: provider runtime images carry their CLI; prove the private-network shape
build-module-runtime.sh adds psql to the postgres image and mc to the minio image
(their clients shell out to those). provider-on-backend-network asserts redis's
runtime, on the backend's private network, binds the broker via NAT and provisions
a consumer with the mesh's credential — the shape the committed provider manifests use.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 00:52:17 +02:00
jschoubben aafa11756a e2e: a provider creates the resource with the mesh's credential (ADR 0053)
Assigns redis as a provider, puts the contributions and unsealed password the mesh
would deliver in its receives path, and authenticates as the consumer with the mesh's
password — PONG proves the login was created with exactly that password (a
self-generated one answers WRONGPASS), with MESH_SEAL_KEY set nowhere.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 00:27:46 +02:00
jschoubben 6067ec1724 e2e: a running runtime picks up a settings change (issue 009)
Assigns grafana configured by settings, changes the token, pushes again, and
asserts the container was replaced (new id) and the rendered config carries the new
value. Builds the host from source, since the behaviour under test is the host's.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 23:40:13 +02:00
jschoubben 0a414d576a e2e: assigned sonarr + grafana prove the two runtime config paths (ADR 0051/0052)
assigned-sonarr proves the Servarr detection path: the runtime discovers its API
key from the app's config.xml and serves its tools. assigned-grafana proves the
settings path: the operator states URL and token as settings, the mesh merges them
into the module's config file, and the runtime serves from that with nothing in the
manifest. Both green.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 23:08:54 +02:00
jschoubben f093769354 e2e: assigned plex + redis prove the module runtime (ADR 0052)
build-module-runtime.sh generalises the audit-logger runtime image to any module
(mesh-tools + sdk + the module's dist, entrypoints for tools/events/provisioner).
Two scenarios and two tests: assigned-plex proves a tools+events module serves its
tools over a mesh-issued scoped account; assigned-redis proves a provider's runtime
serves tools AND runs its provisioner in the same broker-bound process, provisioning
a grant and emitting its lifecycle event. Both green.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 22:29:53 +02:00
jschoubben 37b16cd0e9 events: the assigned audit-logger, proven in the lab (ADR 0048)
test/integration/assigned-audit.test.ts raises a node into a mesh, assigns it
the audit-logger through the control plane, and asserts the mesh delivered a
scoped amqps account (not the broker's own), the host ran the container, and an
emitted event reached the trail — the delivered credential authenticating is
the proof. scenarios/audit-node.yml is the lean single-node bed that stocks the
runtime image. Passes 1/1 against the real lab.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 02:13:01 +02:00
jschoubben ba7f7998f7 events: an e2e test — an emitted event reaches the audit trail over the mesh's broker
Adds test/integration/events.test.ts: raises first-node (which raises the
broker as tier-1 substrate), runs the runtime+audit-logger against that
broker, emits a module and a node event, and asserts they reach the trail
with their metadata read back from ADR 0047 headers (a pure body), plus that
the durable per-consumer queue and mesh.events.dead exchange exist on the
raised broker — asked of the broker, not assumed.

scripts/build-runtime-image.sh builds the self-contained runtime image
(mesh-tools + vendored sdk + audit-logger) it runs, saved to a tar for
MESH_LAB_RUNTIME. The events path itself is verified; the incus raise is the
part a lab run exercises.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 00:56:39 +02:00
jschoubben 01ca02ad04 A machine works through its queue, and the wait allows for it
With caught-up finally an equality, the last red test turned out to be
telling the truth about something real: declarations queue, the machine
applies them one at a time at half a minute each, and by the twenty-
fifth test it is minutes behind the latest push. 240 seconds was not a
generous bound on one apply — it was an accidental bound on the whole
backlog.

Doubled rather than tuned, and the real remedy filed instead: a machine
asked to be five successive things should become the last one, which
is a decision about the link rather than about this timeout
(novox/hq 04-ISSUES/031).
2026-09-02 01:05:51 +02:00
jschoubben 81e5f31910 settled() asks which, not when
The timestamp comparison lost the race between one test's closing push
and the next test's opening one: the old apply's report landed newer
than the new send and settled() passed for a declaration the machine
had not read. The report names its declaration now, the mesh says
whether it is the current one, and this reads the answer instead of
inferring it.
2026-09-02 00:03:00 +02:00
jschoubben a65115fc81 The forge poll covers both halves, and the cache failure names the container
The provisioner makes the role and then the database; a poll that
waited for the first and checked the second once was racing the gap
between two statements, and lost it once, eighteen seconds into a run.

The cache test's provisioner could not resolve the store's name, which
usually means the store's container never registered it — and the
diagnostics showed only the provisioner's side of that conversation. A
grant that never arrives now prints the container states, the store's
log and the provisioner's, so the next failure names the half that
actually fell over.
2026-09-01 23:20:38 +02:00
jschoubben 5e22e6bc49 The resolve-together test carries the whole catalogue
Eleven modules planned as one set on one machine, which is what a real
node looks like. Two are left out by name rather than silently: the
mesh under test already runs a module called registry and an adopted
workload called umami, and adding the catalogue's manifests would
replace the records of things that are live and assigned — the adopted
umami would suddenly require a database it never asked for.
2026-09-01 23:09:49 +02:00
jschoubben 384e0f4e7e Current is not caught up
settled() returned the moment a declaration was current, because the
sent digest is recorded at send — so both new tests asserted on a
machine still applying, and found containers not created yet and
bindings not written. The eternal-waiting fault had been standing in
front of this gap the whole time; fixing it is what let the tests get
far enough to fall in.

Caught up now means the machine's own last report is newer than what
was sent to it — two timestamps the mesh recorded itself, read from the
`reported` section status now carries. The residual latency between
"reported" and "every container answers" stays with the tests' own
polls, where it always was.
2026-09-01 23:03:05 +02:00
jschoubben 2f2f1d931e The cache edge, proven end to end
A consumer contributes a key prefix and gets an ACL user; the test is
that the grant means exactly what the manifest said, in both
directions: its own keys usable, anyone else's refused by the store
itself, and the flush a tenant must never have refused with them.

Waited for through the store rather than through logs: the user list,
asked with the password the host wrote into the server's own conf file
on the machine — nothing invented, both ends reading what the mesh
delivered.

And the forge is asked on the port the mesh assigned, not the one the
module declared. The old curl aimed at 3000, which was right until
ADR 0038 moved the machine side — a latent break that would have fired
on the first run to get past the settling that used to fail first.

The scenario stocks redis and its provisioner, and the rebuild builds
the provisioner image with the others.
2026-09-01 22:45:05 +02:00
jschoubben 0d286b7cc8 What review found in the lab, fixed
A segment named "uplink" is refused. The lab claims that name for the
NAT bridge behind `egress: true`, and a scenario wearing it first would
have its egress machines silently attached to an isolated bridge — a
declared key doing nothing, which is the fault this repo exists to
refuse, in the repo that refuses it.

settled() parses inside the try. A truncated status from a struggling
machine was the one shape of bad answer that still threw out of the
wait, and the likeliest moment for one is exactly the machine the poll
is watching. Malformed now counts as "could not ask", like the exec
that times out.

And a sentence on the uplink's UseDNS saying its inertness is
load-bearing: it matters only where systemd-resolved runs, and on a
machine whose modules own resolv.conf the uplink must not outvote the
resolver a scenario is testing.
2026-09-01 21:54:59 +02:00
jschoubben 4a343a2652 A machine is as big as the scenario says, and may reach the world
Three changes, found by one failing test.

The forge failed three runs in a row as "status hangs", and it was
diagnosed twice as contention — real defects, fixed, and not the cause.
The heartbeats told the truth in the end: every exec on anchor crawled
from 15s to 105s, because eleven containers plus a database pull were
running in a 1GiB machine. Starvation presents as whatever you were
doing when the page-outs start, which is why it wore two other bugs'
clothes first.

So machine size is now the scenario's to declare — memory and cpus per
machine, default unchanged. The anchor that carries the whole substrate
is bigger than the laptop that joins it, and the comment on the
scenario says why in terms of what lands there.

`egress: true` gives a machine one extra interface on a lab-supplied
NAT network, addressed by DHCP because the one address a scenario has
no business choosing is on the host's side of the fence. Declared
per machine and off by default: a closed scenario stays the rule
(novox/hq ADR 0016), and the exception exists because a first node
fetches its images before any mesh can serve them — which is now the
tested path (04-ISSUES/029), and a lab that can never reach upstream
cannot prove the bootstrap it exists to prove. The uplink route is
metric-4096, so it never shadows a route the scenario declared. A
detached machine declaring egress is refused, not ignored.

And settled() treats a poll that threw as a poll that missed. An exec
timeout at minute four of a wait is "could not ask", not a verdict on
the machine.
2026-09-01 21:42:18 +02:00
jschoubben f85dbb0713 The registry is added, not built
A mesh that has just bootstrapped cannot build the module that gives it
an artifact store: building publishes to the store, and the builder will
not start without one (novox/hq 04-ISSUES/029).

This test built it and passed, because the scenario's registry was
already standing to receive the push — which is precisely why a real
first mesh would have hit this and the lab never did. A stand-in for
Docker Hub was quietly also standing in for the thing under test.

So the module now names its image by digest, the way the bundle names
the three a first node starts from, and is added as a manifest rather
than built. That is the only path open to a real first mesh, so it is
the path this walks.

Its skip on MESH_LAB_BUILDER goes with it. Nothing in the test needs a
builder any more, and a skip that names a thing the test does not use
sends the next person to look in the wrong place.
2026-09-01 21:17:59 +02:00
jschoubben 144362be37 A poll that could not ask has not been answered
`settled` used `must`, so a failed exec ended the wait as though the
machine had reported a failure. It had reported nothing: the control
plane is a container on the node being polled, and while that node
applies a declaration an exec into it can lose its stdout fifo to
containerd. The run then blamed the mesh for a question that missed.

Could not ask and asked, and the answer was bad are different facts, and
only the second is the machine's. A failed poll now keeps the reason and
tries again; the timeout reports whichever came last, so a control plane
that is genuinely unreachable still fails the test — with the reason
rather than with a stack trace.

Every five seconds rather than every two. Each poll is an exec into a
container on a machine that is busy applying, and thirty times a minute
was competing with the apply rather than observing it.
2026-09-01 20:25:54 +02:00
jschoubben 2bf5846bf2 Wait for the machine before asking what it is running
The forge test read `status` the instant `push` returned and concluded
the machine was fine. It was describing the apply before this one.

`push` sends and returns — it prints "sent N resource(s)" and the
machine applies afterwards. So every assertion made immediately after
one is racing it, and this race lost quietly: no failure reported, and
a container that did not exist yet read as a container that would never
exist.

`settled` asks the mesh, in its own terms: a node is caught up when it
is neither waiting for what it was sent nor wrong about what it applied
— the two questions `status` already answers, read as JSON so a test is
not parsing a report written for a person. A machine reporting a failure
ends the wait immediately rather than at the timeout, because it will
not become right by being waited for.

The container assertion now also prints the plan. A container missing
because the mesh never asked for it and one missing because the machine
could not make it are one sentence and two entirely different faults,
and the plan is what separates them.
2026-09-01 19:34:23 +02:00
jschoubben e99bffc93e Ask the machine what it did, before asking what it produced
The forge test pushed and then waited for a database login. When the
containers were never created at all, it reported "no login was created"
— true, and silent about why. Two hundred and thirty seconds spent
proving something downstream of the actual failure.

A push being accepted and an apply having worked are different facts,
and this test depends on the second. It now reads what the machine says
about itself, and whether a container exists, before it starts waiting —
and carries the host's own log into the failure either way.

The suite otherwise passed 24 of 25 on this run, which is the first time
the forge reached a clean attempt with nothing upstream blocking it.
2026-09-01 16:44:21 +02:00
jschoubben 6591a2e040 A canary first: one machine, one path, three minutes
Suggested by Jochen, and it paid for itself on its first run.

A suite that takes forty minutes is a suite you hear from once a day.
Every fault found today would have shown up in the first three minutes
of it — a module pinned to an image that does not exist, a consumer
given a password and no name to present with it, a credential file
nothing could read, a search for a password that read the password as an
option. The other thirty-seven minutes proved things that were already
working.

So this runs first, on one machine, with the three images the mesh needs
for itself. It walks one path: a mesh comes up, a module lands, and a
consumer gets a credential it can actually use — the name to present,
the address, the port, and a password only the host could put there.
Deliberately not a smaller copy of the full suite: that path is where
everything went wrong, and a canary checking many things shallowly is a
canary whose failure nobody can read.

`suite` runs it and stops if it dies, saying why rather than leaving
somebody to wonder what the missing thirty-seven minutes would have
said. Skipped when the caller named its own files.

It measured 164 seconds against forty-odd minutes, and failed three
times on its first run for one reason: applying the bundle raises a
control plane but does not tell it a machine exists. I had left out
enrolment, and the long suite would have taken forty minutes to say so.
2026-09-01 16:24:19 +02:00
jschoubben 44d3e53bcb Three failures, all mine, all worth having
**A password beginning with a dash broke the search for it.** The
credential test greps the machine's own files for the delivered
password; this run's password started `-S`, so grep read it as an option
and refused the whole invocation. The test compared the usage message
against "0" and reported the password as leaked. That is the worst way
for a search to fail — it says it found something. Fixed with `-e` and
`--`, which is what those exist for.

**The planning test could not redirect what the scenario does not
serve.** Rewriting an image reference only works for repositories the
scenario's registry actually holds, and the object store's provisioner
was not stocked — so that module kept its placeholder and the refusal
fired, correctly. It is stocked now, so all five are planned again. The
skip path stays for anything genuinely unserved, and says which module
and why: a planning test quietly covering four instead of five is the
false coverage this suite exists to prevent.

**The forge failed because of the one above it.** The planning test
threw before its cleanup could run, leaving a module assigned that
refused the next push, so no database container was ever created.
Yesterday's fix moved that cleanup where a failure cannot skip it — but
`after` still only unassigns what was assigned, so it now tracks what
actually got added rather than what was intended.
2026-09-01 16:16:57 +02:00
jschoubben 2c7e69f4db Plan what could actually run, and put the machine back either way
Two faults, both found by the guard that now refuses a placeholder
digest on its way to a machine.

The planning test added the manifests exactly as they sit on disk, which
includes an image the mesh builds — and that image has no digest until
it is built, so the file legitimately carries a placeholder. The test
was therefore planning something that could never run, which is the
whole complaint. It now points the references at this scenario's
registry first, exactly as the forge test does.

The second is worse and more ordinary. Its cleanup was the last
statement in the test body, so the first failure skipped it and left
five modules assigned. The next test's push was then refused by a module
this one had abandoned — a failure that reads as a fault in the test
that was working. Cleanup that only runs on success is not cleanup, so
it moved to `after`, where a failure cannot skip it.
2026-09-01 15:54:56 +02:00
jschoubben 55022edc1e Run the forge, on a database the mesh gave it
Everything up to now stopped at composing a declaration. That proves the
control plane and the host agree, and proves nothing about whether the
thing described works — which is how five modules sat pinned to images
that did not exist while parsing and resolving perfectly.

The forge is the right one to run first. It needs a database from
another module, a password it did not choose, and a connection string it
could not have written itself: the address and port come from what the
database serves, the user name from what the mesh decided both ends
would call it. If any of that is wrong it cannot start, and nothing else
in this suite would notice.

The test checks the chain in the order it has to happen — the login
exists, the database it owns exists, the forge answers, and its log does
not say authentication failed. That last one matters: a forge that
started and could not reach its database would still answer on its port.

Rewriting image references is now shared rather than copied from the
bundle, which had the same problem first. A digest is not knowable until
something is built, and when it is, it belongs to whichever registry
served it — so the text says which image and the scenario says which
copy. Matching is on the repository, with a test that a repository
ending in another one is not half-replaced.

Also makes the planning test put the machine back. Tests here share one
mesh, so the five modules it assigned were inherited by whatever ran
next; harmless while nothing pushed, and not harmless now.
2026-09-01 15:39:30 +02:00
jschoubben 5137720aa7 A path is not a secret, and the check said it was
The last failure of the run: `MESH_BROKER_FILE=/var/lib/mesh/builder/broker`
reported as "something secret-shaped, which the broker would see".

`/` is in the base64 alphabet, so any absolute path of 24 characters or
more matched the pattern meant to catch a sealed value. An absolute path
is a *reference* to a secret and naming one is the whole design — the
mesh delivers a credential as a file and a module says where.

Excluded explicitly rather than by loosening the pattern, and checked
both ways: a real sealed value and a base64 blob are still flagged, a
relative path still is, only an absolute path is passed over.

Worth the words in the comment. A check that fires on the right shape
for the wrong reason is worse than none — it is the one that gets
suppressed, and then it is not there when it is right.

It surfaced now because tests in this file share one mesh: the builder
was assigned by an earlier test and appears in this one's declaration.
2026-09-01 09:52:28 +02:00
jschoubben 3503ad990b The registry is addressed the way every other machine is
novox/hq 04-ISSUES/024. The registry machine had its address set with
`ip addr add`; every other machine gets a systemd-networkd unit. That
one difference stalled the lab indefinitely.

An address set by hand leaves networkd waiting to configure a link it
was never told about, so the link sits at `configuring` for ever.
`systemd-networkd-wait-online` has TimeoutStartUSec=infinity, so
`network-online.target` is never reached — and Docker is ordered after
it. `docker load` then blocked on a socket whose daemon was queued
behind a target that would never come.

Measured before and after on the same scenario: stuck with five pending
systemd jobs and `docker` inactive; now `enp5s0 configured`, `docker`
active, no jobs, and the whole raise completes in 87.5s.

The guess in the issue was wrong, and it was wrong in the usual way —
stocking had just been changed, so stocking looked guilty. Stocking
takes 34s and always did.

Two things that made this cost hours rather than minutes are fixed with
it. Placing an image now waits for the container runtime to answer and
refuses after 120s naming what systemd is waiting on, so a stall becomes
a failure that says why instead of three stacked timeouts totalling 35
minutes. And the end-to-end test passes `onProgress`, so a raise says
what step it is on — it printed nothing at all until it finished, which
is why 35 minutes of nothing read as a slow test.
2026-09-01 09:37:16 +02:00
jschoubben e83e5ab24e Prove a hole in a config file is filled, on a real machine
Nothing did. The sealed-placeholder substitution and the bound-value
substitution were each covered by unit tests in the repository that
performs them, and the two expressions that find the holes live in
different repositories — so both sides could agree with themselves and
disagree with each other, and the first thing to notice would be a
program connecting to a host called "${bound:postgres-database:at}".

So the consumer in the credential test now ships a configuration file
with four holes in it: the address and port from what the provider
serves, the name to present from what the mesh decided, and the password
sealed. The mesh fills the first three before sending, the host opens
the credential and fills the last on the machine, and the test reads the
file off the machine and checks that the password in it is the same one
the credential file holds — and that no ${ survived.

This is the only place those two mechanisms meet a real host.
2026-09-01 03:15:22 +02:00
jschoubben eb02b8fe6b Assert the whole of what keycloak is given, not one file's name
The credential moved: the sealed password is a password alone, at
`.secret`, and `database.env` is now the connection keycloak could not
have written — address and port from what the provider serves, user name
from what the mesh decided both ends would call this consumer.

So the test asks for both, and for the seam between them: the password
is still a hole, the sealed value travels beside the file that needs it,
and no ${bound:...} survives as a value. That last one matters most —
a placeholder written through would be read as a hostname, and the
failure would name neither the module nor the mesh.
2026-09-01 03:06:59 +02:00
jschoubben cb0edcee8d The end-to-end test reads the grant where the mesh now writes it
Two assertions in the full-mesh test encoded the old naming: the grant
file read back from the provider, and the PostgreSQL role the real
application logs in as. Both are named after the consumer now, and a
consumer is a module on a machine.

These are the two that matter most in this file — it is the only place
where a real application authenticates against a real database with a
password the mesh delivered and cannot read, so they are what would have
caught the naming going wrong end to end.
2026-09-01 02:48:24 +02:00
jschoubben e7f4a49e40 A grant file and a provisioned login name the module too
novox/hq 04-ISSUES/022: a consumer is a module on a machine, not a
machine. The mesh now writes <node>.<module>.secret and the provisioners
name the role and the access key after both.

The fixtures here write what the mesh writes, so they move with it —
that is the whole point of them, and a fixture that kept the old shape
would agree with the bug rather than catch it.

The object-store assertions are the ones that mattered most: one store
holds every bucket behind one endpoint, so isolation is a policy rather
than a property. One access key per machine meant every module on a node
shared it, and the policy confining each consumer to its own bucket
confined none of them.
2026-09-01 02:45:40 +02:00
jschoubben 0ffb24ff5d Write down what a run has to be pointed at
Reconstructed from the source twice now, which is 04-ISSUES/005 in its
own README: a test whose artifact was not pointed at skips rather than
fails, so an unset variable is a green run that proved nothing. The
first attempt today reported "skipped 24" and left a receipt claiming
zero of everything — working exactly as designed, and indistinguishable
at a glance from a suite that had nothing to do.

Also records the two things that cost time either side of it: `check`
says which variables are missing before a long run rather than skipping
quietly, and a heredoc into `newgrp` runs the suite as a child of a
shell that immediately exits, so it needs `setsid nohup … &` or it dies
with the shell that launched it.
2026-09-01 02:44:32 +02:00
jschoubben a516ee847b Adopt a third-party workload, and keep a mesh between runs
**The adoption.** Software nobody here wrote, taking its credentials the
way such software does — from its environment — and needing two
containers that reach each other by name. The first module that could
not have been declared this morning: it needs the network shape and it
needs a sealed value to reach a container's environment.

Its password is accepted rather than generated, which is the whole shape
of an adoption: a service that already exists keeps the credential it
already has. Asserted properly — a wrong password is refused by the same
database, so the passing case means something.

**The warm scenario.** A mesh kept between runs and returned to, which
turned twelve minutes of bootstrap into thirty seconds of restore. Off
unless asked for: a run that is meant to mean something raises from
nothing.

Its guard fired for real during this work, unprompted — a mesh-host
commit landed and it refused the stale base, naming both commits, rather
than passing tests against yesterday's binary. That is 04-ISSUES/005's
rule one level down.

Three things the guard learned the hard way and now handles: a snapshot
captures disk and not memory, so the host is restarted after a restore
and asserted to have come back; the stocked image digests are worked out
while raising and a restored instance never raises, so they are kept;
and comparing only the repositories this run can see clears the ones it
cannot, so both directions are compared.

The one real bug behind five failed attempts was in mesh-host and it
reported itself precisely: a network shape the language had and no host
implemented. Everything else was scaffolding of mine.
2026-09-01 01:20:14 +02:00