Commit Graph
129 Commits
Author SHA1 Message Date
jschoubben 572aae2eb4 A plan with codes, stated up front and reported against
Steps were identified by their own sentences, so 'which one failed' meant reading
prose, and rewording a step silently made it a different step with no history.
Each now carries a stable code: R for raising the mesh, P for it being able to
produce, U for something being used on it, V for verifying what it says about
itself, E for enduring — a change following on its own, and coming back after the
machine stops.

The plan is data, printed before anything is attempted, so a reader knows what
the run intends to establish rather than inferring it from what happens to be
printed. The run ends with a table and a JSON report, and distinguishes SKIP from
FAIL: a step whose dependency failed was never asked, which is not the same as a
step that was asked and said no.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:00:13 +02:00
jschoubben 2f470f27b8 Genesis makes six claims; assert them separately
Genesis is twelve steps and was reported as one line, so a failure said nothing
about which claim broke and a pass was one tick standing in for six things being
true: the substrate up, the control plane built rather than handed over, the
pivot finished, the registry serving what was published into it, the machine
enrolled with an agent actually running, and the builder installed as a module.

Each is asked of the machine rather than read from the installer's own output.
The installer saying it published an image and the registry serving one are
different facts, and only the second matters.

And the catalogue is invoked by its real entrypoint. 'mesh-tools' is not on PATH
in the runtime image; the image runs 'node dist/main.js', and the invoke mode is
missing from the header comment that says there are three modes.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 21:58:58 +02:00
jschoubben 7641bb2059 A one-node mesh, and twelve things that have to be true of it
The common case, and the one that was never tested as a whole. What existed
asked whether four machines converged; it never asked whether ONE machine ends
up holding a mesh.

The order was wrong too. Three machines were enrolled second, into a mesh that
could not yet produce a single module, and that was reported as though something
had been shown. 17-raising-a-mesh is explicit: genesis ends with a mesh that
RUNS, and what remains after the core modules are built is "adding machines".
So the core comes first and machines arrive last — here, not at all, because a
second node is only meaningful once the first is complete.

Three things were missing entirely and nothing complained, because nothing asked:
the mesh never built its own catalogue, never had a store of its own for that
catalogue to use, and never rebuilt its own control plane through the module
path.

And four checks that were absent rather than failing:

  - it can describe itself — status, module list, plan --json, and the
    catalogue's five tools ASKED rather than observed. A container being up was
    being read as the catalogue working, which is the same error as matching a
    container by substring and finding the wrong one.
  - its networking is what the modules asked for — default closed, ssh open,
    declared ports open, .internal names written, module networks present. Left
    out altogether, which is hard to defend given the firewall work this week.
  - a change to a module's source reaches the machine on its own. The capability
    the migration depends on.
  - it comes back after a reboot. Never once tested; the lab had no way to
    restart a machine, because nothing had ever needed one.

Machines are named by role now — anchor, home-server, workstation, laptop — not
after the operator's own nodes, which made test output and real state hard to
tell apart.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 21:52:24 +02:00
jschoubben 9146f30859 Name the container, and issue the account
Waiting for "a container whose name contains lavinmq" was satisfied by the
broker — lavinmq, up and healthy — while the thing under test, the module's own
runtime mesh-lavinmq, crash-looped beside it. The step went green and the fault
was found by reading docker ps by hand. Containers are named exactly now, and a
failure prints that container's own last words.

And lavinmq gets a broker account, which it was never issued. Without one the
mesh still fills the secret the module declares it owns, with a generated value,
so the runtime starts, fails to parse a password as a credential document, and
loops on a JSON syntax error that mentions no missing account.

The two module-issue calls written .catch(() => {}) are not. That pattern has now
hidden three separate faults in this file.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 21:20:57 +02:00
jschoubben c3d5ec1d55 Build each repository from its own ref, and build the provider
lavinmq needs building now, so the step that assigns it builds it first.

And the ref is per repository rather than one value for all of them: a change
under test lives in one repository, and building the others from that branch
would prove it against itself.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 21:05:28 +02:00
jschoubben f04911a763 Give the module the broker it asks for, rather than a module that asks for nothing
The mesh refused to place amqp-ping: nothing provides amqp. That refusal is
right. The substrate raises a broker, but as a bundle resource — plumbing, not a
module the mesh has a record of — so it offers nothing to anything, and a module
wanting a broker wants one in the graph.

lavinmq is that module and needs no building, its image being upstream, so this
is a register and an assign. The alternative was to pick a module with no
requires, which would have passed by testing less.

Also: the control plane's image has no /tmp to copy a manifest into. Root does.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 20:56:35 +02:00
jschoubben 4ef0a19053 A manifest the control plane can open, and stop swallowing the failure when it cannot
mesh-control runs in a container, so a manifest pushed to the machine is not a
file it can read; `module add` said so plainly and it was briefly taken for a
missing manifest. It is copied the last step of the way now.

The base's registration was doing this too, and its failure was swallowed by a
bare catch on the reasoning that the module might already be known. The step
passed regardless — a base with nothing to stand on builds whether or not the
mesh holds a record of it — and the fault surfaced one step later, where the
record was needed.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 20:46:11 +02:00
jschoubben babf08b9f8 Raise machines that are somebody, and a bed that hands over nothing
Every machine in a bed is a clone of one base image, so all of them booted with
the same /etc/machine-id. systemd's DHCP client derives its client identifier
from that file and dnsmasq keys leases on the identifier rather than the MAC, so
four machines with four distinct MACs were handed one address and the host kept
one ARP entry for it. Whichever machine last answered an ARP request received
everybody's replies.

This is the fault behind every run lost to "flaky lab DNS": resolution that works
two times in three, pulls that succeed on a retry, and one machine out of four
being fine while the rest have no path at all. It survived an earlier diagnosis
that blamed resolver ordering, because reordering resolvers on a machine that has
just won the ARP race looks exactly like a fix.

Each machine is now given its own machine-id before the uplink lease is asked
for, and a check after addresses are applied refuses to go on if two machines
took the same one — the positive control this never had, since the fault is
invisible where it happens and unrecognisable where it surfaces.

The egress check also now demands five consecutive lookups rather than one. A
single answer is what let a machine resolving one query in three pass and then
die twenty minutes later inside a pull.

And fresh-mesh: whole-mesh-full's topology with genesis-single's honesty. The
four-machine bed loads thirty-four of the mesh's own images onto its machines
from the workstation because it does not build them, which is a shape no real
installation has and the same fiction the lab removed when it deleted its own
registry. This scenario names no images at all. The machines pull what is public,
the installer builds the control plane, and the mesh builds the rest — including,
last and deliberately, a module on a machine that did not build it.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 20:41:55 +02:00
jschoubben fb6dab6a48 The beds place the builder's manifest too 2026-09-14 12:31:44 +02:00
jschoubben d27f24cf3e The bed resolves an artifact the way the builder would
It pre-builds these images and stocks them, which is the lab standing in for the
builder — so it must do what the builder does and replace the artifact with the
reference the machine holds. Without it the unresolved field travels to the
machine and the whole declaration is refused.
2026-09-13 04:56:00 +02:00
jschoubben fb18807000 The bed checks the control plane was built, not carried
The pivot checks proved the running control plane is pinned to a digest this
mesh's registry serves, which a carried image satisfies just as well. What the
installer now exists to make true is that a build happened, from the commit the
bed asked for — and that was printed and not checked.
2026-09-13 04:28:54 +02:00
jschoubben 607ea241c7 The lab's installer carries a builder, and genesis is told what to build
Both beds now pass a repository and a commit, and refuse to run without them
rather than raising a machine the installer cannot finish.
2026-09-13 04:24:18 +02:00
jschoubben e427e41389 Raise a mesh of one by the installer, in a bed of its own
The four-node bed proves genesis entangled with three machines joining across a
gateway, so the cheapest check of the install path costs a four-machine raise.
This is genesis alone: one machine, the installer, and the question asked of the
machine rather than inferred from an exit code.

Genesis moves into a shared routine both beds call, rather than being described a
second time here — a second description kept in step with the first is what put
the whole procedure inside a fixture to begin with.

The scenario needs two things the first draft missed, and both cost a full raise
to discover: a way out to the internet, because the installer's first act is to
pull the substrate; and a container runtime, because the installer's first refusal
is a machine that has none.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:46:05 +02:00
jschoubben 3f58a0a08f The provider remap moves a port; it does not bind it to loopback
postgres and lavinmq carried 127.0.0.1: in their remap, and it broke a consumer
on a node that has no substrate to collide with. A module is told to reach its
provider at <node>.internal, that name is the node's overlay address, and a
provider listening only on loopback refuses it — letta on ace failed with 'is the
server running on that host and accepting TCP/IP connections?' while postgres sat
healthy beside it.

The collision needed a different port, which is what every other entry here does.
The address was never part of it, and it made the provider unreachable by the one
name the mesh hands its consumers.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 22:05:18 +02:00
jschoubben 555401a787 The bed bootstraps through the installer, not around it
ADR 0067's own acceptance check said the lab must raise its anchor by running the
program a bare machine runs. It did not: whole-mesh-full applied the substrate bundle
by hand and then looped enrolment over all four machines as one continuous operation.
That gets the order right by accident and models the wrong shape — and an install
procedure that exists only as a test fixture is exercised by whoever writes tests and
never by whoever installs, which is why every bootstrap fault this year was found late.

Two acts now, and the first gates the second.

  GENESIS is novox running mesh-bootstrap: the installer is built from source before
  the raise (make bootstrap, carrying the control-plane image built in the same run),
  placed beside the host binary, given the two manifests it reads, and run. The bed
  then asserts a WORKING MESH OF ONE — the control plane answers, the registry replies
  on /v2/, the container called mesh-control is running from a registry-pinned digest
  rather than an image id, the registry agrees it serves it, temp-mesh-control is gone,
  and the mesh has heard from its node. The image-id check is ADR 0067's "the pivot
  completed" verbatim: if it is still an id, nothing was published and this mesh can
  never roll out its own upgrades.

  JOINING is ace, shanks and g14: host binary, token, enrol, run. novox is NOT enrolled
  again — the installer already did it, and a second identity is one the mesh does not
  know.

If genesis stops, the bed prints which of the installer's ten steps it stopped at and
goes no further. A second machine joining a mesh that is not ready is a different
failure, and running it would bury this one underneath it.

The anchor is no longer handed mesh-control:development. Its absence is the point: the
installer carries that image inside itself, and handing it over as well would make the
load say "already held" and leave the carrying untested — the same class of fiction the
lab's own registry used to hide. A unit test asserts the scenario keeps it out.

The registry is reached at 127.0.0.1:5000, which is a finding rather than a shortcut: a
runtime refuses a plain-HTTP registry at any address but a loopback one, so the digest
the control-plane module is pinned to is one only the anchor can pull. Enough here,
because only the anchor runs a control plane. Written down in the bed.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 11:46:42 +02:00
jschoubben 6c09ddb528 An image the machine has no account for is handed over, not fetched
Deleting the lab's registry left the operator's own images to be pulled like
anything else, and they cannot be: their registry wants an account and a
scenario machine has none. The pull fails with 'no basic auth credentials',
which is not something more patience fixes.

So the test is no longer 'did the mesh build it' but 'can the machine get it at
all'. Two ways to fail that — published nowhere, or published somewhere the
machine cannot authenticate to — and one consequence: the workstation, which
does hold the credential, exports it and loads it.

Worth saying what this stands in for. In a finished mesh these are built by the
builder and published to the mesh's own store, and every machine pulls them from
there with a credential the mesh granted. Until that store exists there is
nowhere for them to come from, and handing them over is the closest honest thing
— not a registry the lab invents, which is what was just removed.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 01:11:18 +02:00
jschoubben 94e617915c The home segment moves off 192.168.1.0/24
It is the commonest home LAN range there is, so on an ordinary workstation the
lab's private segment and the machine's own network are the same addresses. The
scenario routes an egress machine explicitly and marks the rest unreachable, so
nothing leaked — but that guard was carrying the whole weight of a collision
nobody chose, and a guard is a bad place for that.

10.99.1.0/24 is still RFC 1918, so the bed still models a home LAN behind an
access point. It is simply far from what this kind of machine already has:
192.168.1 is the LAN, 172.16-31 and 192.168.16-95 are container bridges, and
10.10/10.42/10.208 are a tunnel, the mesh overlay and the virtualisation daemon.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:00:19 +02:00
jschoubben a4c2a9b90b The routing record is 0066, not 0056
0056 was already 'the authority is the control plane, not a database'. The
routing record was renumbered where it lives; these citations pointed at the
wrong decision, which is worse than pointing at none.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:35:35 +02:00
jschoubben 7875145c0e An unpinned tag is a finding, not a reason to stop the bed
Seven catalogue modules — photos, photos-eef, photos-filip, invoicing, novox.be,
de-spiegel, amqp-email-forwarder — name `registry-api.…/novox/…:latest`. That is a TAG,
which ADR 0006 forbids and mesh-host refuses. It has never shown, because the lab's
registry rewrote every reference to a digest it had assigned, tag or not: the fiction
was not only serving the images, it was silently pinning them.

There is nothing to pin them with now. Asserting here would take whole-mesh-full down in
`before()`, before the overlay it exists to prove; the useful outcome is that each of
those modules fails to apply on the node that carries it, saying exactly why, while the
rest of the bed runs. So the reference passes through and the harness says so out loud.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:17:57 +02:00
jschoubben 675facdb0d The beds name images the way a machine would find them
Twenty-eight integration tests each carried their own copy of the same two helpers,
which pointed a manifest and the substrate bundle at whatever the lab's registry had
assigned. They now share two in the harness, and the difference is the point: ours is
rewritten to the ID the machine holds it under, and everything else is left exactly as
written so the machine pulls it.

**The substrate bundle is where the fiction was most load-bearing.** mesh-host's
`examples/substrate-first-node.lock` pins all three of its images at
`192.0.2.250:5000/…`, which is the address the lab's registry served from — it was
written for a target, and the target was the lab. Two of those are ordinary third-party
images and become the digests mesh-catalog's own postgres and lavinmq modules pin, so
the substrate's store and broker are literally the images the mesh runs. mesh-control
exists in no registry at all and becomes the ID the machine was handed. **The bundle
itself should be fixed in mesh-host and this substitution deleted with it.**

Beds that wrote a manifest by hand named an image by repository and let the rewrite
supply a digest. There is nothing to supply one now, so `onTheMachine` refuses an
unpinned reference and hands back the digest the catalogue pins — a bed runs the image
the mesh ships, and a bed that drifts from the catalogue is testing a different
postgres.

Three beds took a third-party image out of the raised list, which no longer contains
one: certificates (pebble), objectstore (minio and its client) and provisioner
(postgres) now name theirs and pull it. builds and mesh publish into the MESH's own
artifact store — the `registry` module's image, on the node, on 5000 — rather than into
scenery the lab raised. That is a different claim, and only one of them exists in
production.

New unit tests cover what a full raise would otherwise be the only way to check: the
routes an egress machine gets (that its gateway is still the path to the rest of the
scenario, that a range with no path is unreachable rather than leaked to the uplink,
that each family gets its own next hop), which machine is handed which of our images,
and the `images:` rule that refuses a third-party entry. The "shipped scenarios are
valid" test now loads every scenario rather than two of them.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:16:41 +02:00
jschoubben 0a0c57b610 whole-mesh-full: the CA root a person hands the mesh must be readable by it
The control plane's image is FROM scratch and runs as 65534, and docker cp keeps
the mode a file had outside — openssl writes a private key 0600 root-owned, so
the copy landed unreadable, secret accept failed with permission denied, and the
CA crash-looped on a root it never got. Chowning it inside the container is not
available: there is no shell in there to do it with.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 22:35:49 +02:00
jschoubben a48b60b604 whole-mesh-full: the bed knows about ADR 0056
The bed set no node a `public-domain` and assigned no `acme-ca` provider, so it
was testing a mesh the design no longer describes — and going green while doing
it, which is the worse half.

**No public domain means no route.** A module now contributes a `label` and
nothing else; the mesh joins it to the node's public domain, and a label with no
domain to join composes to nothing at all. Every routed module on this bed was
therefore unreachable by name, silently, and no assertion noticed. novox now
carries `novox.incus` and ace `zurag.incus` — `.incus`, because this repository's
beds name nothing routable. The workstations carry none, which is also the design
being exercised: a node that does not face outward has no public domain.

**No acme-ca provider means no proxy.** route-proxy requires one, so without a
provider it is unresolvable and takes every routed module with it. step-ca is
assigned on the anchor, at mesh scope, and given an operator root — made with
openssl on the anchor and handed over through the real `secret accept` path,
because the mesh cannot invent a PEM and the random bytes it makes for an
own-secret nobody supplied would leave the CA crash-looping on a root key that is
not a key.

**What is asserted is the half that is decided and cheap**: that each routed
module's name composes to `<label>.<public-domain>` — read from the proxy's own
received-routes file, the mesh's answer on the machine rather than this test's
arithmetic checked against itself — with `@` composing to the bare domain, and
that the proxy answers for one of them over HTTP.

**What is NOT asserted is issuance.** Whether route-proxy obtains a certificate
from step-ca over ACME depends on mesh-control fixes landing as this is written,
and a bed that gated on them would report somebody else's in-flight work as its
own failure. step-ca is listed as a reported gap for the same reason.

The substrate apply also retries up to three times. `raise` now refuses to return
until every machine can fetch a manifest from the scenario registry, so the first
attempt should be the only one; a pull is simply the one step here that can fail
for a reason that goes away by itself, and the cost of not retrying was a whole
raise left as a bare shell.

Typechecks; not run end-to-end — see the ADR 0056 section for what is expected to
fail until the issuance path is fixed.
2026-09-10 21:06:34 +02:00
jschoubben 80b0670ebe whole-mesh-full: the real segmented topology, and the overlay proven across the access point
Rewrite the flat three-node whole-mesh-full (separate anchor, one public segment)
into production's real shape: two segments and one access point. novox sits on
the routable `hosting` segment and IS the anchor — it runs the substrate, its own
service set, the overlay hub and public ingress; there is no separate anchor node.
ace, shanks and g14 sit on the household `home` segment behind a NAT gateway,
reachable from outside only through what they dial out to.

The bed drives, and verifies, the thing the flat beds never could: the WireGuard
overlay forming ACROSS the access point — a home node dialling novox's public hub
endpoint out through the gateway's masquerade, the handshake completing through the
NAT, the keepalive holding the hole open. Phase A proves it (handshake state + a
ping over the overlay) before any heavy module lands; Phase B converges both server
sets. With MESH_LAB_KEEP the instance is raised under a fixed id and left standing.

Collapsing the substrate onto novox exposed real facts the separate-anchor beds
never hit, fixed here:
- the substrate bundle advertises the broker at 192.0.2.10 (the old anchor); a
  token carries that verbatim as the endpoint a node dials, so with the substrate
  on novox it must be novox's own public address. Rewritten at apply (the cert is
  fingerprint-pinned, not hostname-checked, so only the address needs correcting).
- the two provider host-port collisions with the co-located substrate: postgres
  5432 vs the store's 127.0.0.1:5432, lavinmq 5672 vs the broker's 127.0.0.1:5672.
  Both provider host publishes are remapped off the substrate's ports.

And a lab limitation this first large-union bed exposed: the image registry VM took
the profile's default `dir` pool and a ~10GiB root, which the ~28GiB union of both
server sets overflows ("no space left on device"). raiseRegistry now places the
registry on the scenario's copy-on-write pool with a sized (default 80GiB, thin)
root disk, MESH_LAB_REGISTRY_DISK overridable.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 11:17:00 +02:00
jschoubben 591f2a641c whole-mesh-full: prove the dry-run fixes (fail2ban hostable, credential own-secrets)
Re-runs the capstone from main after the dry-run fixes merged.

fail2ban: added to the novox set. The capability fix (intrusion-prevention ->
firewall) makes it HOSTABLE — it is now assigned, not refused — which is the
gate. Its service reaching active is a host concern the offline lab cannot meet
(the VM ships nftables but not fail2ban, and the isolated segment has no route to
the package mirror, so pacman cannot fetch it), so fail2ban joins GAPS_NOVOX: its
failed package resource is tolerated like firewall's oneshot nftables.service.

7 credential sidecars: before the push, a FAKE app credential is delivered for
each (plex/bazarr/ombi/home-assistant/nzbget/qbittorrent on ace, umami on novox)
through the real operator path — `secret accept <node> <module> <name> --from`.
The bed asserts each sidecar advances PAST its old "no credential" crash (it reads
the delivered value); app-auth failure against the real app with a bogus value is
expected and not gated.

Result: SUITE_EXIT=0. Both node-plans converge on one substrate (novox 13/13
core, ace 17/17 core), fail2ban hostable, all 7 sidecars past their crash.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-08 20:05:04 +02:00
jschoubben fcdcd338c0 Add whole-mesh full three-node bed (stage 3: both server sets, one substrate)
Combine the novox (17-module) and ace (24-module) sets on ONE substrate and
prove both node-plans converge together. anchor runs the substrate only;
novox and ace each run their own self-contained set (own postgres/redis), so
nothing crosses a node boundary except enrolment and the shared broker/store.
The four modules both nodes run (postgres, redis, mssql, portainer) are added
once and assigned to each node, each getting its own per-node broker account.
An overlay is placed across all three nodes.

Proven green: both nodes converge together on the one substrate. ace reaches
applied+current with all 17 of its CORE up (and letta too this run); novox
reaches all 13 CORE up with its only failed resource the known firewall.load
oneshot gap. The two node-plans share one broker without collision — distinct
novox-<mod> and ace-<mod> accounts for the modules both run. No new cross-node
bug (overlay/DNS/identity/port) surfaced; ports are per-VM and the sets are
node-self-contained. Tolerates the same nine credential-sidecar gaps and
firewall's nftables.service oneshot documented in the per-server beds.

Resource envelope: 3 VMs (anchor 4GiB, novox 16GiB, ace 18GiB) + registry
scenery, ~79 union images (~35GB) stocked to one registry VM and pulled
concurrently by both nodes; fit within 125GiB host RAM and the 180GiB lab pool.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-08 18:22:56 +02:00
jschoubben 90651e1e73 Add whole-mesh ace dry-run bed (stage 2 of whole-mesh rehearsal)
Install the real ace server's converted service set (24 modules) together on
one node behind the substrate — sibling of the whole-mesh-novox bed, the
media/home-automation half. Loads each committed module.json from
mesh-catalog, rewrites image refs to the scenario registry's digests, remaps
the co-located host-port collisions (qbittorrent/searxng/unifi :8080,
nzbget/unifi :6789), and pre-creates the ADR-0051 operator-owned media
library dirs under /services/media so the media stack's `accesses` resolve.

Proven green: the whole 24-module set RESOLVES and applies (214 resources,
node applied+current) — the ADR-0051 shared-dir `accesses` mechanism works
cleanly across eight co-accessing media modules. The CORE 17 converge whole:
postgres/redis/mssql, sonarr/radarr/lidarr/jackett/tautulli/bookshelf,
mosquitto/influxdb/grafana/baserow/nodered/searxng/unifi/portainer.

Reported as escalated gaps (do not gate green): six tool-runtime sidecars
crash-loop because the committed manifest does not wire the app credential
they need (plex MESH_PLEX_TOKEN, bazarr MESH_BAZARR_API_KEY, nzbget
MESH_NZBGET_URL/PASSWORD, qbittorrent MESH_QBITTORRENT_URL/PASSWORD, ombi
MESH_OMBI_API_KEY, home-assistant MESH_HOMEASSISTANT_TOKEN) — the umami/photos
class from novox; each server is up, only the sidecar is down. sonarr/radarr/
lidarr/jackett/tautulli self-configure from the app's config file and their
runtimes come up. letta's app has a first-boot postgres migration race
(pgvector the deeper blocker, per two-node-db).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-08 00:19:52 +02:00
jschoubben 581fd6da77 Add whole-mesh novox dry-run bed (stage 1 of whole-mesh rehearsal)
Install the real novox server's converted service set together on one node
behind the substrate — the whole-catalogue install this rebuild never ran.
The bed loads each committed module.json from mesh-catalog (no hand-written
manifests), rewrites image refs to the scenario registry's digests, and
remaps the co-located host-port collisions (nextcloud/invoicing/route-proxy
:80, minio/invoicing :9000, gitea/umami :3000).

Proven green: the whole set of 17 modules RESOLVES and applies (191
resources); the CORE 13 converge whole — all five providers (postgres,
redis, minio, mongodb, mssql) plus keycloak, gitea, nextcloud and invoicing
reaching their providers and staying up, plus portainer, verdaccio, registry
and route-proxy.

Reported as escalated gaps (do not gate green): fail2ban (declares
capability intrusion-prevention that no host detector provides, and an
unappliable assignment blocks whole-node resolution), umami/photos/mailu
(catalog manifests do not wire the runtime/app env the images need; photos'
server image is an alpine placeholder), and firewall (nftables.service is a
oneshot that exits, but the module declares state running so mesh-host marks
it failed).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 22:54:10 +02:00
jschoubben cf26b19b96 local-model bed: prove model-access answered by a node (ADR 0055)
A single-node VM bed: an ollama provider and a local-model consumer are
assigned; the resolver answers the consumer's model-access with the local node
(no licence demanded), the consumer's openai.env is templated with the served
endpoint, and a request to it reaches the running model server. Proves the
node-answer of model-access end to end (ollama on host network, keyless).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 05:25:21 +02:00
jschoubben a1231a7f89 openai bed: prove the static-key model-access path (ADR 0050)
A single-node VM bed: the operator sets an API key on an openai licence, the
mesh seals it to the consumer, the host unseals and mounts it, and the consumer
writes it as OPENAI_API_KEY (env + Codex auth.json). Asserts the written key
equals the one set — the other shape ADR 0050 defines, and the ADR 0054 branch
where a static-key vendor records no usage. Green on the first run.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 04:25:26 +02:00
jschoubben 9cc3c1ac2a model-usage bed: prove the usage store end to end (ADR 0054)
A VM bed: a postgres provider and the model-usage consumer on one node, the
substrate on the other. A usage event injected into the mesh is upserted into
model-usage's provisioned store, asserted at both grains, latest-per-key, and
in the clear.

Also, in build-module-runtime.sh, add migrate/index.ts and pg.d.ts to the
compiled entrypoint set so a module may carry a run-once entry and an ambient
type declaration (model-usage uses the latter for the pg driver).

The bed surfaced and drove several fixes elsewhere: a short module slug for the
S3-key identity bound (ADR 0049), host-network containers getting the mesh's
names (mesh-control), and injecting the event from a publisher rather than the
pure-consumer store (its account has no publish right by design).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 04:07:41 +02:00
jschoubben 87c4820130 anthropic bed: package a module's own npm deps, stage grant files readably
Two harness fixes the green end-to-end run needed:

- build-module-runtime.sh installs a module's non-@novox runtime deps under
  /app/modules/<module>/node_modules, so a module can carry a private dependency
  (the anthropic-manager seals with tweetnacl-sealedbox-js). The shared tree still
  answers @novox/* and common packages. A no-op for modules that declare none.

- stageIntoControl chmods the manager's 0600 adopt/refresh outputs to 0644 on the
  anchor host before docker cp, so the distroless mesh-control (non-root, no chmod)
  can read the staged file. What is staged is a sealed box or the access token,
  never a cleartext refresh token.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 02:47:50 +02:00
jschoubben 71bea08f3b anthropic-bed: prove the host unseals the refresh token, no node-key stub
The bed follows the reworked flow: the manager module seals the refresh token to the node's
PUBLIC key, the HOST unseals it and mounts the cleartext at the manager's bound path, and the
refresh reads that cleartext -- no fake node key pair is mounted any more, the host uses its
own real sealing key.

  - the manager is a model-access holder deployed first, so its bound facts (carrying the node
    public key) are delivered; the consumer is added only once an access token exists to seal.
  - adopt reads the node public key from the bound facts; the test asserts the host mounts the
    cleartext refresh token for the manager, and that it reaches nowhere on the consuming node.
  - the refresh_grant assertion reads { sealed, manager_key }.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 01:55:28 +02:00
jschoubben 6be5072565 anthropic-bed: prove model-access refreshes on the manager node
A lab bed for Phase C of model-access (ADR 0050), OAuth endpoint stubbed.
It drives the real runtime images through the whole flow: the manager
seals a refresh token at rest and opens it on the manager node alone,
mesh-control is handed only the access token and an opaque re-sealed
envelope via licence submit-refresh, and the consumer writes an
access-token-only credential. Asserts the refresh token -- original and
rotated -- is nowhere on the consuming node and only ciphertext in the
control plane's database.

build-module-runtime.sh also compiles adopt/refresh/apply/usage
entrypoints. Stubbed and flagged: the vendor endpoint, the manager node's
private key (mounted; a host capability to deliver it does not exist
today), and the submit transport (the test invokes the CLI on the
manager's output).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 01:00:37 +02:00
jschoubben f04c4ba3be lavinmq-bed: GREEN — AMQP provider provisions a require-only consumer's vhost (mint fix)
Two-node: substrate broker on anchor, lavinmq provider + amqp-ping consumer on laptop. The
consumer gets its scoped vhost+user, connects, and round-trips a message. Requires the
mesh-control require-only-mint fix. Diagnostic removed now it's green. SUITE_EXIT=0.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 23:20:59 +02:00
jschoubben 73bd086193 lavinmq-bed: reproduces a credential-mint gap for require-only consumers of a parameterless provision
The lavinmq AMQP provider comes up and serves, but its receives file has given:[] — the
consumer amqp-ping (requires amqp, contributes nothing, as a parameterless provision like
redis-cache takes no per-consumer payload) is never minted a credential, so the provisioner
creates no vhost. Diagnostic in the test dumps the empty grants file + the provisioner log.
This is the resolve/plan mint path (grantsFor -> SecretsFrom), ADR 0048 territory, and it
likely affects redis-cache consumers the same way. Preserved for a focused fix; not merged.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 16:09:01 +02:00
jschoubben 8e494c0e78 Add lavinmq-bed: a two-node amqp provider + consumer bed
A lavinmq provider and an amqp-ping consumer ride laptop while the
substrate's own broker owns 5672 on anchor — the twin of two-node-db.
lavinmq is the mesh's control broker AND a user-facing capability, so a
provider must publish 5672 for its consumers and cannot share a node
with the control broker that already owns it; the split unblocks the
chain single-node.

The bed proves, layered: the run-once bootstrap computed the admin hash
and wrote the broker config before the broker started (ADR 0052); the
service and both runtimes are up and stable; each module got its scoped
broker account on the substrate broker; the provisioner created the
consumer's vhost AND user, both named for the derived login; and the
consumer connected to that vhost with the mesh-minted password and
round-tripped a message. The consumer uses ${bound:amqp:as} for user
and vhost both, and the provider's serves carries the port.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 15:50:33 +02:00
jschoubben 6cd595d403 route-forwarding: GREEN — slug hello-web so it resolves; Host-routing + withdrawal proven
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 15:16:56 +02:00
jschoubben f03647d605 Add route-forwarding lab bed proving the route grant end to end
A scenario and integration test assign route-proxy (provider) and hello-web
(consumer) on one node, then assert a request to the consumer's name -- sent to
the proxy -- is forwarded to the workload and returns its answer, and that
unassigning the consumer withdraws the route so the same request stops working
(the proxy replaces its table rather than merging). Modeled on
mesh-grant-end-to-end and schedule-tick: module add, assign, one push, settled,
with no module issue (route-proxy needs no scoped account).

build-route-proxy-image.sh compiles the Go proxy from
mesh-control/examples/route-proxy into mesh-route-proxy:development for the
scenario to stock. This bed proves route-forwarding over plain HTTP;
public-ACME TLS is proven separately by certificates.test.ts against a real
ACME server.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 15:07:23 +02:00
jschoubben 76ecbaed85 lab: a scheduled container fires on its cadence without gating (ADR 0053)
The bed that proves the scheduled-container primitive end to end. schedtest
is the thinnest carrier of ADR 0053: one container marked
schedule: "* * * * *" that appends a timestamp to a mounted data dir each
time the host fires it -- no service, no listener, no provisioner, no
runtime, no tools, no events.

The three claims it proves, from the ADR's "How each claim is checked":
installing the schedule leaves the node current WITHOUT a run (baseline
captured right after settled, the deliberate inversion of run-once); the
container fires on its cadence (a line beyond the baseline within ~150s);
and it recurs (a second line on the next minute -- cadence, not a one-shot).

schedtest serves and consumes nothing and carries no runtime, so it is not
issued a broker account: module add -> assign -> one push is the whole
sequence, no module issue. The tick image is a bare alpine served by the
scenario's registry by digest.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 14:20:54 +02:00
jschoubben a9ecce25cb Two-node DB-consumer bed + a scenario disk field
The GREEN multi-node regression bed that proves the DB-consumer gate: substrate/control on
one node, postgres+redis providers and baserow+letta consumers on another, each consumer
getting its own credential and its own mesh-named database across the overlay. Requires the
mesh-control provider-seal-key fix and the mesh-catalog db-name fix.

Includes a general lab capability: a machine 'disk' field sizing the VM root disk (a broad
install exhausts the pool default and the host fails mid-apply with 'no space left on
device'). The bed sets 60GiB.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 13:48:26 +02:00
jschoubben e703a94fcd tools-confluence: the tools-only bed for confluence (serves 3 tools, no creds)
Sixth green regression bed. Same shape as tools-gitlab: a runtime-only module comes
up under the mesh, serves its full tool surface with no valid credentials (the Servarr
lesson), stays up, binds its serve queues, and gets its scoped broker account.
SUITE_EXIT=0.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 01:54:21 +02:00
jschoubben c9619a17f4 Add tools-gitlab lab scenario proving the tools-only install
Prove gitlab — the exemplar tools-only, outbound-only external-SaaS
integration — installs: a scenario assigning gitlab to one node, and a test
asserting the mesh-runtime-gitlab container comes up and stays up, logs
[mesh-tools] serving 23 tool(s), binds its serve queues on the broker, and
gets its own scoped account — all with NO valid GitLab token, the case the
Servarr lesson is about.

No gitlab arm is needed in build-module-runtime.sh: gitlab speaks HTTP and
needs no extra CLI in the image.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 01:33:52 +02:00
jschoubben 62e479b9fe catalogue-mqtt: prove the run-once primitive end to end
A bed that assigns mosquitto and asserts the run-once step seeded dynsec
before the broker: the bootstrap ran to completion (not left running), the
seed is on disk owned by the broker's uid, the broker is up and stable
(it crash-loops against an unseeded store, so a stable broker is the proof),
and the node reached current. On top, the seeded admin authenticates over
MQTT and the provisioner grants a scoped client a consumer connects with.

build-module-runtime.sh gains a mosquitto arm (install mosquitto_ctrl from
the mosquitto package — it is not in mosquitto-clients on bookworm, and a
musl binary from eclipse-mosquitto would not load) and compiles the module's
bootstrap/index.ts entrypoint.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 00:33:41 +02:00
jschoubben fc36fb5b51 catalogue-media: two modules accessing one operator-owned directory co-resolve (ADR 0051)
sonarr and radarr both access /services/media/downloads — the exact duplicate path
the resolver refused before novox/hq ADR 0051 (04-ISSUES/036, 012). Each now declares
it as an `access`, not a `directory` resource, so the pair co-resolves and one push
configures both. The operator provides the shared media dirs before apply (the host
refuses an absent access); the bed creates them after enrol and before the push.

Proves: the push is not refused, the node converges once, both modules' server and
runtime containers are up, and both server containers mount the same operator-owned
spool.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 23:02:41 +02:00
jschoubben 571a6cd6fd WIP: catalogue-apps install bed (mongodb, unifi, marrytts)
Adds the mongodb runtime CLI (mongosh) to build-module-runtime.sh, a
catalogue-apps scenario, and its install test. Proven so far: the ADR-0054 slug
applies and the mesh accepts the push (mongodb consumer identity mesh_anchor_mongo
fits). NOT green: the node applies but never reaches 'current' within 1200s — a
persistent reconcile divergence (applied-but-never-current, no crash), likely a
module declaring a resource its container mutates (issue-011 class). Needs live
VM inspection to name the module. Not merged.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 15:30:50 +02:00
jschoubben ab04dd814f Add catalogue-small co-residence bed (four modules, one push)
Raises the first-node substrate and assigns postgres, minio, redis and
plex to one anchor in a single push, proving they resolve and come up
together on one node. postgres and minio each get a consumer that
connects with a real granted credential.

redis follows the corrected provider contract (ADR 0048, issue 032): its
runtime reconciles the contributions the mesh delivers at MESH_RECEIVES
and creates each consumer's ACL user with the mesh-minted password,
sealing nothing — no MESH_SEAL_KEY, no *.grant.json/*.credential path.
The provisioning proof authenticates as the consumer with the mesh's
password (PONG), matching the green provider-uses-mesh-credential bed.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 14:14:56 +02:00
jschoubben cbd9b647ec e2e: unblock the minio grant — the consumer declares a slug (ADR 0054)
Issue 010 fixed: bucketuser declares slug `bkt`, so its identity mesh_anchor_bkt (15)
fits an S3 access key where mesh_anchor_bucketuser (22) did not. Unskips the test.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 02:51:53 +02:00
jschoubben b7ba7534af e2e: the whole grant for an S3 bucket (skipped — blocked on hq issue 010)
Mirrors the postgres bed for minio: a provider (runtime carries mc) + a consumer
requiring s3-bucket, proving the consumer reaches its bucket with the access key and
secret the mesh delivered. It surfaced a real limit: the mesh derives `as` =
mesh_<node>_<module> (22 chars), and an S3 access key is capped at 20, so minio refuses
the service account. The test is correct and skipped pending 04-ISSUES/010, not worked
around.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 01:58:01 +02:00
jschoubben de7f3b9c05 e2e: the whole grant for a database — a consumer connects with what the mesh delivered
Assigns a postgres provider (its runtime carries psql) and a module that requires
postgres-database; the mesh mints one password, postgres's provisioner creates a role
and database under the mesh's login with it, and the consumer connects to its database
with the delivered credential (a password-checked connection) — select 1. Nothing placed
by the test. The postgres half of the per-backend provider proof (ADR 0052/0053).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 01:47:57 +02:00
jschoubben 617b5577f7 e2e: the whole grant, mesh-driven — a consumer authenticates with what the mesh delivered
Assigns a redis provider and a module that requires redis-cache; the mesh mints one
password, seals a copy to each end, writes redis its contributions and the consumer
its bound file, and the host unseals each side. redis's provisioner creates the ACL
user under the mesh's login with the mesh's password, and the consumer's delivered
credential authenticates (PONG). Nothing is placed by the test — the provider/consumer
contract (ADR 0053) working as one thing, no shared key anywhere.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 01:14:48 +02:00