Commit Graph
95 Commits
Author SHA1 Message Date
jschoubben 290be37a93 step-ca: mount the directory, not each secret file
A file bind mount tracks the inode. The host writes atomically — new file, rename
over — so the container keeps reading the file that was there when it started,
and a rotated or newly-delivered secret never reaches it. Mounting the parent
directory resolves the path on each open instead.

This is a known shape (hal KB troubleshooting/docker-bind-mounts), and it cost an
hour here before it was looked up: the CA crash-looped on a root key it had
already been given, because the container still held the inode from before the
key arrived.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 08:49:55 +02:00
jschoubben 0d2c2dd989 minio: pin the manifest these machines can actually run
The pinned digest named the arm64 manifest, so an x86 machine pulled it and the
container died with 'exec format error' on every restart. It never showed while
the lab ran its own registry: the harness pushed the WORKSTATION's copy, which
is amd64, and every machine then pulled that under a digest the registry had
just assigned. The registry was quietly correcting the architecture too.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 02:51:34 +02:00
jschoubben 9ff320ecca Pin the operator's own images by digest, as every other image already is
Nine references across seven modules named ':latest'. ADR 0006 forbids it and
the host refuses it by name — and the refusal had never fired, because the lab
pushed every image into its own registry and rewrote each reference to the
digest it had just assigned. Deleting that registry made these the only
manifests the host would now reject (novox/hq 04-ISSUES/039).

The digests are what each tag resolves to today, read from the registry that
serves them.

This is a stopgap and should be said as one: a digest written into a repository
is wrong the moment anybody rebuilds, which is precisely why the design has the
repository name artifacts and the mesh hold digests. Until something builds and
publishes, a digest that is stale is still better than a tag that silently moves.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 01:06:14 +02:00
jschoubben 4a1c5ac817 The artifact store requires nothing, because at genesis there is nothing
A route-label migration gave the registry a public name, and with it a
requirement. But the registry is the first module a new mesh installs: at that
moment nothing provides a route, so the push is refused and a mesh cannot get
its own image store — the cycle 04-ISSUES/029 closed, re-entered through a
different door.

The same rule that record states applies: a module providing the artifact store
cannot depend on what the store is needed to deliver. A public name for it is an
ordinary want and belongs to a module beside it, installed once there is a mesh
to install things.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:06:23 +02:00
jschoubben 71f8012c5b The control plane as an ordinary module
novox/hq ADR 0067 pivots genesis through a temporary control plane and then
reinstalls the control plane as an ordinary module pinned to a digest the mesh's
own registry assigned. That record notes the one thing missing: a control-plane
module manifest, which did not exist.

It could not be written honestly before now. The control plane read its store
connection from MESH_STORE_<CONTEXT>, that connection string carries a password,
and a manifest can put a sealed value into a file's `content` but has nothing
that substitutes into a container's `env`. So the manifest could carry the
password in the clear, or omit the setting. mesh-control now also accepts
MESH_STORE_<CONTEXT>_FILE, which is how every other module here is given secret
material, and the manifest follows.

What the substrate bundle gives the control-plane container today, and where
each part has gone:

  MESH_STORE_INVENTORY   own-secret `inventory`, mounted, named by _FILE
  MESH_STORE_IDENTITY    own-secret `identity`,  mounted, named by _FILE
  MESH_STORE_LICENCES    own-secret `licences`,  mounted, named by _FILE
  MESH_BROKER_AMQP       own-secret `broker`, through an env-file hole
  MESH_BROKER_MANAGEMENT own-secret `broker-management`, likewise
  MESH_BROKER_ADDRESS    ${machine:at}:5671 in that same env-file
  MESH_BROKER_CERTIFICATE  plain env; the path is not a secret
  network host, args ["serve"], the broker's TLS volume  unchanged

The two broker URLs go through an env-file rather than a file of their own
because mesh-control has no MESH_BROKER_AMQP_FILE. That is the same fault one
layer over, and the same remedy would fix it; it is out of this change's scope
and is written down rather than papered over.

None of these values is in the manifest. Each is an own-secret the operator
supplies with `secret accept` — the mesh cannot invent a connection string — and
the container restarts when any of them changes.

The module claims `the-control-plane` at mesh scope, which the bundle has no way
to say: two control planes writing one inventory is a fault worth refusing at
assignment. It carries no `listens`, because `serve` dials the broker and binds
nothing. The image is the catalogue's placeholder digest for a mesh-built image,
which the installer replaces with what the registry assigned.

Checked with the real parser: all 67 manifests through catalogue.ParseManifest
and every module's CheckIdentity against all four node names — 0 problems — and
this manifest rendered through Resolution.Declaration, so the ${secret:…} names,
${machine:at}, the restart-on ids and the image pin are exercised rather than
merely parsed. Slug `control`: mesh_shanks_control is 19 of the 20 an S3 access
key keeps.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:43:07 +02:00
jschoubben 16b4dc9ab9 step-ca: certify the machine the mesh reaches it at
Its API certificate carried localhost only, so a proxy dialling the address the
mesh handed over refused it on hostname verification. The names now compose from
the machine the module was assigned to, which a manifest could not know and now
does not have to (mesh-control ${machine:...}).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:12:48 +02:00
jschoubben d9fbc72565 step-ca: give the CA's own secret material to the uid the CA runs as
The upstream `smallstep/step-ca` image runs as uid 1000. An own-secret lands as
a file the host writes root:root 0600 — the module says only a name and a path,
so there is nowhere to say who must be able to read it — and the container
crash-looped on `permission denied` reading its own root key. The lab got past
it first with `chmod 0644`, which hands the root key to every local user, and
then by running the CA as root, which is worse.

Neither is needed. A module composes its own files, and a `file` resource takes
both a `mode` and an `owner`, with `${secret:name}` reaching the module's own
secrets — the mechanism redis already uses to hand its password to a server
running as 999. So the three pieces of init material are declared as owned
files: still 0600, owned by 1000:1000, and those are what the container mounts.

The raw own-secret files stay where they were. Nothing mounts them now; they are
how the secret comes to exist on the machine, and the host is the only thing
that reads them. Same shape as redis's `default.secret`.

Verified against the real code rather than by inspection: mesh-control attaches
the sealed value to each of the three resources and keeps `owner` and `mode`
(`sealedFor` + `intoFile` over this manifest), and mesh-host parses `owner` on a
sealed-substituted file and lands it 0600 owned by 1000:1000 (`applyFile`).
2026-09-10 20:57:58 +02:00
jschoubben 43ca9c9c73 Give every module that would overflow its login a short slug
An identity is `mesh_<node>_<slug-or-name>` and a backend keeps 20 characters
(an S3 access key). Overflow makes a module unresolvable, and this catalogue
was finding it one module at a time, on a raise: route-proxy on novox is 22,
home-assistant on ace is 23. Two found by hand where a sweep would have found
eighteen.

So the whole catalogue was swept instead, against the longest node name the
mesh actually has (`shanks`, six characters) rather than against the node each
module happens to sit on today — a module is assigned somewhere, and where is
not a property of the manifest. That leaves eight characters for the identity
source, and eighteen modules were over it.

Slugs added, chosen to stay greppable in a provider's user list:

  anthropic-consumer  claude     openai-consumer     openai
  anthropic-manager   anthmgr    portainer           portain
  audit-logger        audit      public-acme         pubacme
  bookshelf           books      qbittorrent         qbt
  cloudflare-dns      cfdns      resolv-conf         resolv
  confluence          confl      resolved-split-dns  splitdns
  home-assistant      hass       route-proxy         rproxy
  invoicing           invoice    verdaccio           verdacc
  mosquitto           mosq
  nextcloud           ncloud

A slug changes the login the mesh mints, so a module already provisioned under
its full name is re-minted under the slug and its old login withdrawn — which
is the provisioner's ordinary business, but it is a change, not a no-op.

Checked with the real parser: every one of the 66 manifests through
`catalogue.ParseManifest`, and every module's `CheckIdentity` against all four
node names. 0 problems, where the same check over the parent commit reports 44.
2026-09-10 20:57:45 +02:00
jschoubben 08be6683b6 route labels: nextcloud is 'drive' (real hostname), novox.be is the apex '@'
nextcloud's label was migrated from a wrong module-name default; its real
production hostname is drive.novox.be. novox.be is the bare-domain apex, now the
'@' label (composeName gained apex support).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 00:27:25 +02:00
jschoubben f0665ba956 ADR 0056: selectable ACME issuer, step-ca init from operator root, route labels
Piece A + B of ADR 0056, completing the internal-CA work in ff01ada.

Selectable issuer. acme-ca is now a role two providers can satisfy: step-ca
(internal CA) or the new public-acme (a fact-only module, no container/listen)
that serves Let's Encrypt production. A mesh assigns one or the other to satisfy
route-proxy's `requires: acme-ca`.

One directory shape for both. A provider serves the ACME directory's parts the
way the mesh already models any reachable service -- an address (`at`), a `port`
and a `path` -- and route-proxy composes `https://<at>:<port><path>`. step-ca
lets the mesh fill `at` (its node) and `port` (its single listen) and serves only
`path`; public-acme, not being a mesh service, serves all three (overriding `at`
with the public host). Same composition either way.

Empty root means the system trust store. Both providers serve `root`: step-ca
the operator root PEM (settled per mesh), public-acme an empty string. route-proxy
writes it to the CA bundle file unconditionally; the binary now reads an empty
bundle as "the root is already trusted by the OS" and falls back to system roots
(examples/route-proxy/main.go, committed on the mesh-control ADR-0056 branch).

step-ca inits from the operator's root. The operator's root cert, root key and
root-key password are mounted at the smallstep entrypoint's default init paths
(/run/secrets/root_ca.crt, root_ca_key, root_ca_key_password) with the matching
DOCKER_STEPCA_INIT_*_FILE vars, so `step ca init` adopts the operator's root
instead of self-generating one -- the CA that signs is the CA route-proxy trusts.

Route names are labels, not FQDNs. Every routed module now contributes a `label`
(the leftmost subdomain) instead of a full public hostname; the node's public
domain composes the name. Apex (novox.be) is left as a full name -- composeName
has no empty-label/apex convention yet (mesh-control follow-up).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 00:21:42 +02:00
jschoubben ff01adaa45 Add internal ACME CA (step-ca) as a generic acme-ca provision; wire route-proxy to consume it
Piece C of ADR 0056: an internal authority certifies routed names by the same
path a public one would, with the proxy pointed at whichever issuer the mesh
names and trusting that issuer's root.

step-ca (new): provides acme-ca (mesh scope), listens 9000 from mesh, persists
its CA home under uid 1000. The root CA cert reaches consumers as a served
value settled from a per-mesh operator setting (no baked root); the init
password and root key are sealed secrets, not literals. No route-name -> IP
hosts map (that is Piece B). Upstream smallstep/step-ca pinned by Docker Hub
digest.

route-proxy: requires + binds acme-ca. ACME_DIRECTORY is composed on the
consumer side from ${bound:acme-ca:at}:${bound:acme-ca:port}, and ACME_CA_BUNDLE
is a file whose content is ${bound:acme-ca:root} -- both from the binding,
replacing the hardcoded directory and the baked root PEM.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 23:07:54 +02:00
jschoubben d018f1e16e Route names mirror real production hostnames from live Traefik
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 14:04:17 +02:00
jschoubben e26ca38eaa novox conversions: give the web apps public names through route-proxy
Add a `route` contribution (requires/contributes/binds) to every web app so
each gets a Host-routed public name via route-proxy, mirroring the
de-spiegel/only-office pattern:

- novox: gitea, keycloak, nextcloud, umami, invoicing, verdaccio, registry,
  and mailu (single mail.novox.be -> 7080; admin/webmail/api ride that port).
- ace: grafana, sonarr, radarr, lidarr, bazarr, ombi, tautulli, jackett,
  nodered, searxng, home-assistant, bookshelf, baserow.

Split photos so its three sites each get a name: photos keeps server +
admin-client (photos.novox.be), and new photos-eef (eef.novox.be) and
photos-filip (filip.novox.be) modules carry the client sites.

The production FQDN stays the literal default; a per-node .incus name is a
settings override applied where the mesh runs, not a manifest hardcoding.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 12:51:15 +02:00
jschoubben 431310fb03 novox conversions: short slugs for the login-length cap + real roundcube digest
The whole-mesh-novox lab validation found three modules unresolvable: a
provision-consuming module's minted login is mesh_<node>_<module>, and on novox
`mesh_novox_only_office`(22)/`_de_spiegel`(21)/`_amqp_email_forwarder`(31)
overflow the 20-char S3 access-key cap (ADR 0049). Added a short `slug` each
(office/spiegel/emailfwd → 17/18/19 chars). Also fixed mailu's roundcube image:
the pinned digest returns "manifest unknown" from ghcr; corrected to the real
:1.9 digest.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 02:35:47 +02:00
jschoubben b3309a0ce9 mailu: full multi-container stack from the live novox deployment
Rebuild the mailu manifest from the running production deployment as the
source of truth, and finish wiring its tool runtime.

- Full 11-container topology: front, smtp, imap, admin, antispam, antivirus,
  webmail, webdav, fetchmail, resolver, redis — real ghcr.io/mailu images
  pinned by digest at tag 1.9 (clamav/radicale/fetchmail digests newly fetched).
- Admin DB now consumes the mesh postgres-database provider (requires +
  contributes + binds + secrets, DB_* templated from ${bound}/${secret}),
  replacing the bundled postgres:13 admindb the live stack still runs.
- Full config env from the live containers as a plain env-file; SECRET_KEY,
  the admin API token and the initial-admin password become own-secrets;
  DB password comes from the provider secret. No secret values hardcoded.
- Enable the Mailu admin REST API (API=true, WEB_API, API_TOKEN own-secret) so
  the ported tools can reach it — the live deployment runs this API OFF.
- mesh-mailu tool runtime on the mailu network: admin API over the module
  network, token from the mounted own-secret, docker.sock for the doveadm mail
  reads, broker + mergeable config with restart-on.
- client.ts fromEnv reads the API token from its mounted own-secret file
  (MESH_MAILU_API_KEY_FILE), matching the cloudflare-dns/umami pattern.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 01:04:55 +02:00
jschoubben 8f99b5bd42 Add amqp-email-forwarder daemon module
Owned daemon that consumes from the mesh amqp provision (lavinmq) and
forwards to SMTP. No ports, no route. AMQP host/port/user/password
templated from the amqp grant (analog: amqp-ping); vhost/exchange/queue
are the app's own config; SMTP user+password are given own-secrets
(secret accept). See report for the EMAILDELIVERY_T vhost integration flag.
2026-09-09 00:41:06 +02:00
jschoubben c423e21a8e Add novox.be; replace photos immich stub with the real app
- novox.be: owned www:latest container, public route novox.be
  (host 4000 -> app 8080). No DB, no secrets. Analog: de-spiegel/hello-web.
- photos: the existing module.json was a wrong immich stub (alpine image,
  port 2283). Replace it with the user's real photo app: photos-server
  backend consuming the mesh s3-bucket (bucket photos) + mongodb-database
  (db photos) providers, env templated from the provisions (analog:
  invoicing), plus the three static client containers (admin + two family
  sites). Only photos.novox.be is routed: contributes.route is single-valued
  across the catalog, so eef/filip need their own modules (flagged).
2026-09-09 00:36:48 +02:00
jschoubben 1fc10be11b Convert only-office, n8n, de-spiegel to nox modules
Mirror the proven converted modules: pure manifests, no runtime.

- only-office: owned onlyoffice/documentserver container, generated JWT
  own-secret, public route office.novox.be. Analog: gitea/hello-web.
- n8n: consumes mesh postgres-database + redis-cache providers instead of
  bundling them, owned container, public route n8n.novox.be. Analog:
  baserow (postgres+redis) + hello-web (route).
- de-spiegel: owned container built from its own repo image, public route
  de-spiegel.novox.be, given SMTP own-secrets. Analog: hello-web/invoicing.
2026-09-09 00:30:06 +02:00
jschoubben 6a84e97e53 Merge pull request 'module fixes from the whole-mesh dry-run: fail2ban capability + tool-runtime credential wiring' (#20) from feat/module-cred-fixes into main 2026-09-08 18:45:24 +02:00
jschoubben 5973d41966 umami: read the admin password from its mounted secret file
The whole-mesh dry-run found umami's runtime crash-looping "admin password is
not set": its `admin` own-secret is mounted at /run/secrets/admin, but the
client read the bare env UMAMI_ADMIN_PASSWORD, which nothing sets. Same shape as
the six tool-runtime credential fixes — read the mounted file first
(MESH_UMAMI_ADMIN_PASSWORD_FILE), falling back to the env. (photos and mailu
remain deeper conversion jobs — a stub app image and a full Mailu config env —
not credential-wiring, tracked separately.)

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-08 18:43:48 +02:00
jschoubben d289e5a928 modules: wire tool-runtime app credentials as own-secrets
The six modules that run a mesh-<mod> tool-runtime sidecar read an app
credential from an env var the manifest never provided, so the sidecar
crash-looped in the whole-mesh dry-run (e.g. "no Plex token — set
MESH_PLEX_TOKEN"). These are operator-set app secrets, so deliver them the
same way cloudflare-dns delivers its API token: an own-secret file mounted
read-only, with a MESH_<APP>_*_FILE env pointing at the mount, and the
runtime code preferring that file (falling back to the existing env so
nothing regresses).

- plex: own-secret token -> /run/secrets/token, MESH_PLEX_TOKEN_FILE
- bazarr: own-secret api-key -> /run/secrets/api-key, MESH_BAZARR_API_KEY_FILE
- ombi: own-secret api-key -> /run/secrets/api-key, MESH_OMBI_API_KEY_FILE
- home-assistant: own-secret token -> /run/secrets/token, MESH_HOMEASSISTANT_TOKEN_FILE
- nzbget: own-secret password -> /run/secrets/password, MESH_NZBGET_PASSWORD_FILE (URL stays plain env)
- qbittorrent: own-secret password -> /run/secrets/password, MESH_QBITTORRENT_PASSWORD_FILE (URL stays plain env)

The operator now completes each with `secret accept <node> <module> <name> --from <file>`.
tsc passes for all six.
2026-09-08 18:40:23 +02:00
jschoubben 75fb16bbfb fail2ban: require the firewall capability, not the non-existent intrusion-prevention
The whole-mesh dry-run found fail2ban unassignable on every node: it declared
`capabilities: ["intrusion-prevention"]`, which mesh-host has no detector for
(its detectors are container-runtime, package-manager, service-manager,
firewall, overlay, graphical-session, seat, privileged). intrusion-prevention
is what fail2ban PROVIDES, not a host capability it needs. It bans via
iptables/ufw, so it needs `firewall` — the same capability the firewall module
declares. The `the-intrusion-prevention` claim (node-exclusive) is unchanged.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-08 18:27:34 +02:00
jschoubben e0456746c3 Merge pull request 'local model: an ollama provider + a local-model consumer (ADR 0055)' (#19) from feat/local-model into main 2026-09-07 05:26:02 +02:00
jschoubben 29fe7c5bd0 local model: an ollama provider + a local-model consumer (ADR 0055)
Model access answered by a NODE, not a licence. `ollama` runs the model server
on host network (0.0.0.0:11434) and provides model-access at node scope, serving
its port and model — it mints nothing, so it is server-only, no runtime. The
`local-model-consumer` requires model-access, gets the endpoint (no secret), and
its templated openai.env carries OPENAI_BASE_URL=http://<at>:<port>/v1 +
OPENAI_MODEL. One provision, two answers: the mesh's own model behind the same
interface as a vendor's.

Proven end to end by the mesh-lab local-model bed (green).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 05:25:21 +02:00
jschoubben b33f9b14c7 Merge pull request 'openai-consumer: the static-key model-access consumer (ADR 0050)' (#18) from feat/openai-access into main 2026-09-07 04:25:56 +02:00
jschoubben c454c6a54c openai-consumer: the static-key model-access consumer (ADR 0050)
The consumer half of the OTHER model-access shape. Where anthropic-consumer
receives a refreshed access token, this receives one operator-supplied API key
the mesh sealed to it and the host unsealed at its secret path — no manager, no
refresh, no usage. It writes the key where an OpenAI/Codex client reads it: an
OPENAI_API_KEY env file and the publicly-known Codex auth.json. Pure node, no
SDK import — the simplest a model-access consumer gets.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 04:25:26 +02:00
jschoubben 4cc1b976f5 Merge pull request 'model-usage: the vendor-neutral usage store (ADR 0054)' (#17) from feat/usage-store into main 2026-09-07 04:08:21 +02:00
jschoubben e8c3ea8979 model-usage: the vendor-neutral usage store (ADR 0054)
The home ADR 0050 left open for a usage reading. mesh-control is a CLI and
cannot consume events, so the store that keeps the current usage picture is a
MODULE — the audit-logger's sibling: it consumes `module.*.usage.*` and upserts
each reading into its own provisioned postgres store, latest per
(licence, consumer, period, metric), in the clear. One vendor-neutral table
holds BOTH grains; they differ only in `consumer` (the holding module for the
licence grain, the session for the finer one). The consumer creates its table
on startup and, as a restart-until-ready service, self-heals rather than
gating the apply on a run-once that must reach a provider over the overlay.

The vendor->row normalisation moves into the adapter, as ADR 0054 requires:
anthropic-manager (licence grain, utilization%) and anthropic-consumer (session
grain, token/cost) now emit already-normalised { rows: UsageRow[], raw } on
their existing keys, so the store stays vendor-blind.

Proven end to end by the mesh-lab model-usage bed (green): a usage event
emitted into the mesh is upserted at both grains, latest-per-key, in the clear.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 04:07:26 +02:00
jschoubben bfbdf90496 Merge pull request 'anthropic model-access: manager + consumer modules' (#16) from feat/anthropic-module into main 2026-09-07 02:48:51 +02:00
jschoubben 19666ff054 anthropic-manager: seal over audited tweetnacl-sealedbox-js, not a hand-transcribed NaCl
The manager reseals a rotated refresh token to the node key with crypto_box_seal. That
seal was a full inline transcription of TweetNaCl's XSalsa20-Poly1305 and blakejs' BLAKE2b
(dependency-free, ~440 lines). Replace the internals with the audited
tweetnacl-sealedbox-js library — the same crypto_box_seal, on the same tweetnacl and blakejs
the mesh used to validate the seal during Phase C.

The exported API is unchanged: seal(value, recipientPublicB64) -> base64. The wire format is
unchanged too — ephemeralPub(32) followed by the box, nonce = blake2b(ephemeralPub +
recipientPub, 24) — so the host's Go box.OpenAnonymous still opens it. The mesh-control
cross-check fixture is regenerated from this seal().

The library and tweetnacl are added to the module's package.json dependencies so the runtime
image bundles them (blakejs arrives transitively). A local ambient .d.ts types the untyped
CJS bundle; it is imported as a default import because Node's ESM loader cannot see a UMD
bundle's named exports.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 02:17:37 +02:00
jschoubben 4c98bee043 anthropic-manager: seal the refresh token to the node key, do no crypto to open
The manager module drops its bespoke ECIES at-rest envelope and the node-private-key mount.
A module is never given a node's private key, so it cannot open an envelope -- the refresh
token is now delivered to it as cleartext by the host, unsealed from an ordinary sealed box.

  - sealedbox.ts: a dependency-free NaCl crypto_box_seal (node:crypto for X25519, transcribed
    XSalsa20-Poly1305 and BLAKE2b-24), byte-compatible with Go's box.SealAnonymous. It SEALS
    only -- opening is the host's job. Proven by a cross-language test in mesh-control.
  - adopt: reads the node's PUBLIC key from the delivered bound facts and seals the operator's
    refresh token to it, handing out only the box.
  - refresh: reads the refresh token as cleartext the host mounted, calls the vendor, re-seals
    a rotated token to the node's public key, submits only { access token, box }.
  - module.json: a model-access holder now -- binds the facts, binds the refresh token as a
    sealed secret; no keys dir, no MESH_NODE_SEALING_* mount.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 01:55:19 +02:00
jschoubben c206e2e11e anthropic model-access modules: manager (refreshable-grant) and consumer
Phase C of vendor-agnostic model-access (ADR 0050/0054). Two TypeScript
runtime modules:

- anthropic-manager: the refresh token is sealed at rest to the manager
  node's own key (atrest.ts, envelope encryption over X25519) and opened
  ONLY on the manager node. adopt seals the first envelope; refresh opens
  it, calls the Anthropic OAuth token endpoint, re-seals a rotated refresh
  token, and hands the control plane only the access token plus the opaque
  envelope. Also polls licence-grain usage (ADR 0054).
- anthropic-consumer: writes the delivered access token to
  ~/.claude/.credentials.json, access-token-only, atomically (the refresh
  token is never delivered); reports session-grain usage from the CLI
  transcripts; a fail-closed identity guard (expected-uuid plumbing is a
  flagged TODO).

Both run as scheduled containers (ADR 0053). Pure logic covered by
node --test fixtures (at-rest round-trip, credential strip, transcript
sum, refresh merge).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 01:00:12 +02:00
jschoubben cd46d53464 Merge pull request 'lavinmq: a user-facing AMQP provider (vhost-per-login)' (#15) from feat/convert-lavinmq into main 2026-09-06 23:21:27 +02:00
jschoubben 97707480d6 Add lavinmq amqp provider and amqp-ping consumer
lavinmq becomes a provider of a user-facing amqp interface: a consumer
that requires a message queue is given its OWN broker — a scoped vhost
and user on a lavinmq provider — not an account on the mesh's own
control-plane broker (ADR 0048). Vhost-per-login is the isolation model,
the exact analog of postgres's database-per-login: the provider names a
vhost after the consumer's login and a user with full rights on that
vhost and none elsewhere, so a login is a broker the consumer alone can
reach.

The provider drives lavinmq through its HTTP management API (client.ts,
the module's one impure seam), with a run-once bootstrap that computes
the RabbitMQ-compatible password hash lavinmq's config wants from the
plain admin secret the mesh mints — the value no ${secret:...}
placeholder can produce and the reason the bootstrap exists (ADR 0052).
serves.amqp carries the port so consumers reference ${bound:amqp:port}.

amqp-ping is a demo consumer: it contributes nothing (the vhost is the
login), reads its grant from an env-file the mesh fills, and uses
${bound:amqp:as} for BOTH its username and its vhost — the db-name
lesson applied to AMQP. It speaks AMQP 0-9-1 over a raw socket with no
npm dependency (the way redis speaks RESP) and round-trips one message.
It carries a slug so its identity fits the 20-char backend bound
(ADR 0049).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 15:50:19 +02:00
jschoubben 090a839d42 Merge pull request 'route-proxy: the native ingress module (drop traefik)' (#14) from feat/route-proxy-module into main 2026-09-06 15:17:21 +02:00
jschoubben 89254ef2c8 hello-web: add slug so its identity fits the 20-char backend bound (ADR 0049)
mesh_anchor_hello_web is 21 chars, one over the S3 access-key bound; slug 'hello' brings it to 17.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 15:16:56 +02:00
jschoubben 36640f68e6 Add route-proxy provider and hello-web consumer modules
route-proxy is the shipping form of the reference reverse proxy (novox/hq
ADR 0007, 08-connectivity section 3): it provides route, is given every
consumer as the file at receives.route, and forwards by the Host header. It
ships the Go proxy from mesh-control/examples/route-proxy via a multi-stage
Dockerfile; no broker, own-secret or provisioner, since it only reads the file
the mesh writes.

ACME_DIRECTORY defaults to Let's Encrypt staging and is overridable per node to
production, so there is no hardcoded production default -- resolving novox/hq
04-ISSUES/004. hello-web is a minimal consumer that requires route and
contributes name+port, to exercise the grant.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 15:07:16 +02:00
jschoubben 9dbbfbb5b4 Merge pull request 'fix: postgres consumers connect to the mesh-named db (the login), not a hardcoded name' (#13) from fix/consumer-db-name into main 2026-09-06 13:47:34 +02:00
jschoubben d301e05794 fix: postgres consumers connect to the db the mesh named (the login), not a hardcoded app name
The postgres provisioner creates each consumer a database named after the login the mesh
minted (mesh_<node>_<module>), per ADR 0048 -- 'a same-named database under exactly that
login'. But baserow, letta, umami, gitea, keycloak and nextcloud each hardcoded their app db
name (DATABASE_NAME=baserow, /letta, /umami, NAME=gitea, /keycloak, POSTGRES_DB=nextcloud),
so the service connected to a database that does not exist ('database letta does not exist').
Each now uses ${bound:postgres-database:as} as the db name -- the login, which is also the db
name -- matching the working meshboard pattern and the provisioner's actual behaviour.

Found by the two-node DB-consumer lab install (mesh-lab assigned-two-node-db); baserow is
proven connecting and running there. The S3 bucket name is the same class of assumption and is
a separate follow-up (s3 identity has its own length bound).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 13:47:16 +02:00
jschoubben dc12df18b4 Merge pull request 'fix: DB/cache/store providers must serve their port (migration blocker)' (#12) from fix/db-provider-serves-port into main 2026-09-06 03:31:50 +02:00
jschoubben 1f0c3ac69e fix: db/cache/store providers must serve their port; drop baserow's empty redis contribution
The breadth install (catalogue-broad) surfaced two latent resolution bugs that
single-module compile checks never caught (compile != resolve):

1. postgres/redis/mongodb/minio served no "port", yet seven consumers
   (baserow, letta, invoicing, gitea, umami, keycloak, nextcloud) reference
   ${bound:<provision>:port}. A binding auto-carries at/from/as; the port is the
   provider's half of the answer and must be declared in `serves` (manifest.go:
   "Serves is what a consumer needs to know ... a port, a path, a realm"). Added
   port to each: postgres 5432, redis 6379, mongodb 27017, minio 9000. Without it
   no DB consumer could resolve, let alone deploy.

2. baserow declared an empty contribution `contributes: {"redis-cache": {}}`.
   redis's serves spec for the cache is empty (a cache takes no per-consumer
   payload), so a consumer only `requires` it; an empty contribution is refused.
   Dropped it (requires/binds/secrets unchanged).

Found by mesh-lab catalogue-broad; the same fixes are mirrored in that bed.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 02:38:22 +02:00
jschoubben dba4554002 Merge pull request 'Convert confluence + jira on the tools-only template' (#11) from feat/convert-atlassian into main 2026-09-06 01:54:51 +02:00
jschoubben db8b60f4e4 Convert confluence and jira into nox catalog modules
Port the HAL confluence and jira integrations to the tools-only,
outbound-only external-SaaS pattern proven by the merged gitlab module:
runtime-only containers (container-runtime capability), own-secret token +
broker, a settings-managed config.json for public config, and a
per-module runtime image.

Each module carries its own Atlassian API client and tools (ADR 0039),
translated from HAL's @hal/sdk zod-schema/MCP-content shape into
mesh-sdk's input/run-returns-data shape. Public config (ATLASSIAN_URL,
ATLASSIAN_EMAIL) lives in config.json; the API token is the one
own-secret. Clients are built lazily and never throw at registration, so
each runtime serves its full tool surface with no credentials (the
Servarr lesson) — confluence serves 3 tools, jira serves 8.

jira's periodic ticket-poller (update-tickets.service/.timer) is NOT
ported: the mesh has no scheduled-task primitive yet (a pending
decision). Only jira's tools are ported; a top-of-file note records the
deferral.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 01:49:22 +02:00
jschoubben 2c3f241010 Merge pull request 'Convert gitlab — the tools-only external-SaaS exemplar (23 tools, lab-proven)' (#10) from feat/convert-gitlab into main 2026-09-06 01:40:31 +02:00
jschoubben 7eb156a82a Convert gitlab into a tools-only mesh catalog module
Port the HAL gitlab module (whose tools lived in @hal/sdk) into a
self-contained mesh-catalog module modelled on cloudflare-dns: the GitLab
API client and all its tools live in the module (ADR 0039), served through
mesh-sdk's registerModuleTools harness.

Tools-only, outbound-only external-SaaS shape: a runtime-only container on
network:host, no service, no listener, no provisioner. The token is an
own-secret; GITLAB_URL is a public setting in a merge:json config file.

The client is built lazily and never throws at registration, so the runtime
comes up and serves all 23 tools even with no valid token (the Servarr
lesson) — it only fails when a tool is actually invoked unconfigured.

Ported 23 tools: projects (list, get), merge requests (list, get, create,
approve, add note), pipelines (list, get, retry, cancel, list jobs, job log),
and project + group CI/CD variables (list, get, create, update, delete each).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 01:33:43 +02:00
jschoubben 09b140fe86 Merge pull request 'Convert baserow, letta, invoicing (DB/S3 consumers)' (#9) from feat/land-db-consumers into main 2026-09-06 01:14:30 +02:00
jschoubben 8b82e47889 invoicing: point app/api at the real registry host
The conversion left the image host as a placeholder (registry-api.example).
Point both containers at the actual registry the built app/api images live in
(novox/invoicing-{app,api}, tags 51-55 + latest present). Tracks :latest to
match the app's continuous-deploy model; pin by digest later if reproducibility
of a specific build is wanted.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 01:14:08 +02:00
jschoubben 515b1e79c7 Convert baserow, letta, invoicing to mesh catalog DB consumers
Mirror the postgres-consumer shape (keycloak/nextcloud): requires the
backing service(s), binds/secrets for the delivered credential, and a
server container that reads it from an interpolated env file.

- baserow: consumes postgres-database + redis-cache; tooled (dormant
  until an account is configured, like gitea's token).
- letta: consumes postgres-database; tooled, live via a mesh-minted
  server-password injected into both server and runtime.
- invoicing: consumes mongodb-database + s3-bucket; plain two-container
  service (app + api), no tools.

Digests pinned for baserow and letta; invoicing keeps private-registry
tags (DIGEST-UNRESOLVED). redis-cache and mongodb-database consumer
shapes are inferred (no prior consumer in the catalog).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 01:12:50 +02:00
jschoubben 3b5f214f3d Merge pull request 'mosquitto: run-once dynsec bootstrap + fix the exit-0-on-failure provisioner' (#8) from feat/mosquitto-bootstrap into main 2026-09-06 01:06:08 +02:00
jschoubben e58cbc1c46 mosquitto: don't read a rejected dynsec command as success
mosquitto_ctrl's dynsec subcommands exit 0 even when they fail — a
"Client not found", an "already exists", a "Connection error: Not
authorized", an "Unable to connect" all return status 0 and report the
failure only as a line of text (verified live against 2.0.11). ctl()
trusted the exit code, so clientExists()'s getClient probe never threw
and always returned true; createScopedClient therefore took the
setClientPassword branch, never ran createClient, and the consumer's
client never landed in the dynsec store — while the provisioner logged
it as provisioned. That is the assigned-catalogue-mqtt failure.

ctl() now scans the combined stdout/stderr for the tool's error markers
and raises a match as the failure it is. Surfacing those errors exposed
two calls that only "worked" by being swallowed: addRoleACL re-run
reports "already exists" (now ignored like createRole), and addClientRole
re-run reports a bare "Internal error" that cannot be told from a real
fault — so the binding is checked with clientHasRole and only added when
absent. Fresh provision, idempotent re-provision, password rotation, bad
admin auth and unreachable broker all verified against a live broker.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 00:57:45 +02:00