Commit Graph
90 Commits
Author SHA1 Message Date
jschoubben 16b4dc9ab9 step-ca: certify the machine the mesh reaches it at
Its API certificate carried localhost only, so a proxy dialling the address the
mesh handed over refused it on hostname verification. The names now compose from
the machine the module was assigned to, which a manifest could not know and now
does not have to (mesh-control ${machine:...}).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:12:48 +02:00
jschoubben d9fbc72565 step-ca: give the CA's own secret material to the uid the CA runs as
The upstream `smallstep/step-ca` image runs as uid 1000. An own-secret lands as
a file the host writes root:root 0600 — the module says only a name and a path,
so there is nowhere to say who must be able to read it — and the container
crash-looped on `permission denied` reading its own root key. The lab got past
it first with `chmod 0644`, which hands the root key to every local user, and
then by running the CA as root, which is worse.

Neither is needed. A module composes its own files, and a `file` resource takes
both a `mode` and an `owner`, with `${secret:name}` reaching the module's own
secrets — the mechanism redis already uses to hand its password to a server
running as 999. So the three pieces of init material are declared as owned
files: still 0600, owned by 1000:1000, and those are what the container mounts.

The raw own-secret files stay where they were. Nothing mounts them now; they are
how the secret comes to exist on the machine, and the host is the only thing
that reads them. Same shape as redis's `default.secret`.

Verified against the real code rather than by inspection: mesh-control attaches
the sealed value to each of the three resources and keeps `owner` and `mode`
(`sealedFor` + `intoFile` over this manifest), and mesh-host parses `owner` on a
sealed-substituted file and lands it 0600 owned by 1000:1000 (`applyFile`).
2026-09-10 20:57:58 +02:00
jschoubben 43ca9c9c73 Give every module that would overflow its login a short slug
An identity is `mesh_<node>_<slug-or-name>` and a backend keeps 20 characters
(an S3 access key). Overflow makes a module unresolvable, and this catalogue
was finding it one module at a time, on a raise: route-proxy on novox is 22,
home-assistant on ace is 23. Two found by hand where a sweep would have found
eighteen.

So the whole catalogue was swept instead, against the longest node name the
mesh actually has (`shanks`, six characters) rather than against the node each
module happens to sit on today — a module is assigned somewhere, and where is
not a property of the manifest. That leaves eight characters for the identity
source, and eighteen modules were over it.

Slugs added, chosen to stay greppable in a provider's user list:

  anthropic-consumer  claude     openai-consumer     openai
  anthropic-manager   anthmgr    portainer           portain
  audit-logger        audit      public-acme         pubacme
  bookshelf           books      qbittorrent         qbt
  cloudflare-dns      cfdns      resolv-conf         resolv
  confluence          confl      resolved-split-dns  splitdns
  home-assistant      hass       route-proxy         rproxy
  invoicing           invoice    verdaccio           verdacc
  mosquitto           mosq
  nextcloud           ncloud

A slug changes the login the mesh mints, so a module already provisioned under
its full name is re-minted under the slug and its old login withdrawn — which
is the provisioner's ordinary business, but it is a change, not a no-op.

Checked with the real parser: every one of the 66 manifests through
`catalogue.ParseManifest`, and every module's `CheckIdentity` against all four
node names. 0 problems, where the same check over the parent commit reports 44.
2026-09-10 20:57:45 +02:00
jschoubben 08be6683b6 route labels: nextcloud is 'drive' (real hostname), novox.be is the apex '@'
nextcloud's label was migrated from a wrong module-name default; its real
production hostname is drive.novox.be. novox.be is the bare-domain apex, now the
'@' label (composeName gained apex support).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 00:27:25 +02:00
jschoubben f0665ba956 ADR 0056: selectable ACME issuer, step-ca init from operator root, route labels
Piece A + B of ADR 0056, completing the internal-CA work in ff01ada.

Selectable issuer. acme-ca is now a role two providers can satisfy: step-ca
(internal CA) or the new public-acme (a fact-only module, no container/listen)
that serves Let's Encrypt production. A mesh assigns one or the other to satisfy
route-proxy's `requires: acme-ca`.

One directory shape for both. A provider serves the ACME directory's parts the
way the mesh already models any reachable service -- an address (`at`), a `port`
and a `path` -- and route-proxy composes `https://<at>:<port><path>`. step-ca
lets the mesh fill `at` (its node) and `port` (its single listen) and serves only
`path`; public-acme, not being a mesh service, serves all three (overriding `at`
with the public host). Same composition either way.

Empty root means the system trust store. Both providers serve `root`: step-ca
the operator root PEM (settled per mesh), public-acme an empty string. route-proxy
writes it to the CA bundle file unconditionally; the binary now reads an empty
bundle as "the root is already trusted by the OS" and falls back to system roots
(examples/route-proxy/main.go, committed on the mesh-control ADR-0056 branch).

step-ca inits from the operator's root. The operator's root cert, root key and
root-key password are mounted at the smallstep entrypoint's default init paths
(/run/secrets/root_ca.crt, root_ca_key, root_ca_key_password) with the matching
DOCKER_STEPCA_INIT_*_FILE vars, so `step ca init` adopts the operator's root
instead of self-generating one -- the CA that signs is the CA route-proxy trusts.

Route names are labels, not FQDNs. Every routed module now contributes a `label`
(the leftmost subdomain) instead of a full public hostname; the node's public
domain composes the name. Apex (novox.be) is left as a full name -- composeName
has no empty-label/apex convention yet (mesh-control follow-up).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 00:21:42 +02:00
jschoubben ff01adaa45 Add internal ACME CA (step-ca) as a generic acme-ca provision; wire route-proxy to consume it
Piece C of ADR 0056: an internal authority certifies routed names by the same
path a public one would, with the proxy pointed at whichever issuer the mesh
names and trusting that issuer's root.

step-ca (new): provides acme-ca (mesh scope), listens 9000 from mesh, persists
its CA home under uid 1000. The root CA cert reaches consumers as a served
value settled from a per-mesh operator setting (no baked root); the init
password and root key are sealed secrets, not literals. No route-name -> IP
hosts map (that is Piece B). Upstream smallstep/step-ca pinned by Docker Hub
digest.

route-proxy: requires + binds acme-ca. ACME_DIRECTORY is composed on the
consumer side from ${bound:acme-ca:at}:${bound:acme-ca:port}, and ACME_CA_BUNDLE
is a file whose content is ${bound:acme-ca:root} -- both from the binding,
replacing the hardcoded directory and the baked root PEM.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 23:07:54 +02:00
jschoubben d018f1e16e Route names mirror real production hostnames from live Traefik
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 14:04:17 +02:00
jschoubben e26ca38eaa novox conversions: give the web apps public names through route-proxy
Add a `route` contribution (requires/contributes/binds) to every web app so
each gets a Host-routed public name via route-proxy, mirroring the
de-spiegel/only-office pattern:

- novox: gitea, keycloak, nextcloud, umami, invoicing, verdaccio, registry,
  and mailu (single mail.novox.be -> 7080; admin/webmail/api ride that port).
- ace: grafana, sonarr, radarr, lidarr, bazarr, ombi, tautulli, jackett,
  nodered, searxng, home-assistant, bookshelf, baserow.

Split photos so its three sites each get a name: photos keeps server +
admin-client (photos.novox.be), and new photos-eef (eef.novox.be) and
photos-filip (filip.novox.be) modules carry the client sites.

The production FQDN stays the literal default; a per-node .incus name is a
settings override applied where the mesh runs, not a manifest hardcoding.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 12:51:15 +02:00
jschoubben 431310fb03 novox conversions: short slugs for the login-length cap + real roundcube digest
The whole-mesh-novox lab validation found three modules unresolvable: a
provision-consuming module's minted login is mesh_<node>_<module>, and on novox
`mesh_novox_only_office`(22)/`_de_spiegel`(21)/`_amqp_email_forwarder`(31)
overflow the 20-char S3 access-key cap (ADR 0049). Added a short `slug` each
(office/spiegel/emailfwd → 17/18/19 chars). Also fixed mailu's roundcube image:
the pinned digest returns "manifest unknown" from ghcr; corrected to the real
:1.9 digest.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 02:35:47 +02:00
jschoubben b3309a0ce9 mailu: full multi-container stack from the live novox deployment
Rebuild the mailu manifest from the running production deployment as the
source of truth, and finish wiring its tool runtime.

- Full 11-container topology: front, smtp, imap, admin, antispam, antivirus,
  webmail, webdav, fetchmail, resolver, redis — real ghcr.io/mailu images
  pinned by digest at tag 1.9 (clamav/radicale/fetchmail digests newly fetched).
- Admin DB now consumes the mesh postgres-database provider (requires +
  contributes + binds + secrets, DB_* templated from ${bound}/${secret}),
  replacing the bundled postgres:13 admindb the live stack still runs.
- Full config env from the live containers as a plain env-file; SECRET_KEY,
  the admin API token and the initial-admin password become own-secrets;
  DB password comes from the provider secret. No secret values hardcoded.
- Enable the Mailu admin REST API (API=true, WEB_API, API_TOKEN own-secret) so
  the ported tools can reach it — the live deployment runs this API OFF.
- mesh-mailu tool runtime on the mailu network: admin API over the module
  network, token from the mounted own-secret, docker.sock for the doveadm mail
  reads, broker + mergeable config with restart-on.
- client.ts fromEnv reads the API token from its mounted own-secret file
  (MESH_MAILU_API_KEY_FILE), matching the cloudflare-dns/umami pattern.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 01:04:55 +02:00
jschoubben 8f99b5bd42 Add amqp-email-forwarder daemon module
Owned daemon that consumes from the mesh amqp provision (lavinmq) and
forwards to SMTP. No ports, no route. AMQP host/port/user/password
templated from the amqp grant (analog: amqp-ping); vhost/exchange/queue
are the app's own config; SMTP user+password are given own-secrets
(secret accept). See report for the EMAILDELIVERY_T vhost integration flag.
2026-09-09 00:41:06 +02:00
jschoubben c423e21a8e Add novox.be; replace photos immich stub with the real app
- novox.be: owned www:latest container, public route novox.be
  (host 4000 -> app 8080). No DB, no secrets. Analog: de-spiegel/hello-web.
- photos: the existing module.json was a wrong immich stub (alpine image,
  port 2283). Replace it with the user's real photo app: photos-server
  backend consuming the mesh s3-bucket (bucket photos) + mongodb-database
  (db photos) providers, env templated from the provisions (analog:
  invoicing), plus the three static client containers (admin + two family
  sites). Only photos.novox.be is routed: contributes.route is single-valued
  across the catalog, so eef/filip need their own modules (flagged).
2026-09-09 00:36:48 +02:00
jschoubben 1fc10be11b Convert only-office, n8n, de-spiegel to nox modules
Mirror the proven converted modules: pure manifests, no runtime.

- only-office: owned onlyoffice/documentserver container, generated JWT
  own-secret, public route office.novox.be. Analog: gitea/hello-web.
- n8n: consumes mesh postgres-database + redis-cache providers instead of
  bundling them, owned container, public route n8n.novox.be. Analog:
  baserow (postgres+redis) + hello-web (route).
- de-spiegel: owned container built from its own repo image, public route
  de-spiegel.novox.be, given SMTP own-secrets. Analog: hello-web/invoicing.
2026-09-09 00:30:06 +02:00
jschoubben 6a84e97e53 Merge pull request 'module fixes from the whole-mesh dry-run: fail2ban capability + tool-runtime credential wiring' (#20) from feat/module-cred-fixes into main 2026-09-08 18:45:24 +02:00
jschoubben 5973d41966 umami: read the admin password from its mounted secret file
The whole-mesh dry-run found umami's runtime crash-looping "admin password is
not set": its `admin` own-secret is mounted at /run/secrets/admin, but the
client read the bare env UMAMI_ADMIN_PASSWORD, which nothing sets. Same shape as
the six tool-runtime credential fixes — read the mounted file first
(MESH_UMAMI_ADMIN_PASSWORD_FILE), falling back to the env. (photos and mailu
remain deeper conversion jobs — a stub app image and a full Mailu config env —
not credential-wiring, tracked separately.)

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-08 18:43:48 +02:00
jschoubben d289e5a928 modules: wire tool-runtime app credentials as own-secrets
The six modules that run a mesh-<mod> tool-runtime sidecar read an app
credential from an env var the manifest never provided, so the sidecar
crash-looped in the whole-mesh dry-run (e.g. "no Plex token — set
MESH_PLEX_TOKEN"). These are operator-set app secrets, so deliver them the
same way cloudflare-dns delivers its API token: an own-secret file mounted
read-only, with a MESH_<APP>_*_FILE env pointing at the mount, and the
runtime code preferring that file (falling back to the existing env so
nothing regresses).

- plex: own-secret token -> /run/secrets/token, MESH_PLEX_TOKEN_FILE
- bazarr: own-secret api-key -> /run/secrets/api-key, MESH_BAZARR_API_KEY_FILE
- ombi: own-secret api-key -> /run/secrets/api-key, MESH_OMBI_API_KEY_FILE
- home-assistant: own-secret token -> /run/secrets/token, MESH_HOMEASSISTANT_TOKEN_FILE
- nzbget: own-secret password -> /run/secrets/password, MESH_NZBGET_PASSWORD_FILE (URL stays plain env)
- qbittorrent: own-secret password -> /run/secrets/password, MESH_QBITTORRENT_PASSWORD_FILE (URL stays plain env)

The operator now completes each with `secret accept <node> <module> <name> --from <file>`.
tsc passes for all six.
2026-09-08 18:40:23 +02:00
jschoubben 75fb16bbfb fail2ban: require the firewall capability, not the non-existent intrusion-prevention
The whole-mesh dry-run found fail2ban unassignable on every node: it declared
`capabilities: ["intrusion-prevention"]`, which mesh-host has no detector for
(its detectors are container-runtime, package-manager, service-manager,
firewall, overlay, graphical-session, seat, privileged). intrusion-prevention
is what fail2ban PROVIDES, not a host capability it needs. It bans via
iptables/ufw, so it needs `firewall` — the same capability the firewall module
declares. The `the-intrusion-prevention` claim (node-exclusive) is unchanged.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-08 18:27:34 +02:00
jschoubben e0456746c3 Merge pull request 'local model: an ollama provider + a local-model consumer (ADR 0055)' (#19) from feat/local-model into main 2026-09-07 05:26:02 +02:00
jschoubben 29fe7c5bd0 local model: an ollama provider + a local-model consumer (ADR 0055)
Model access answered by a NODE, not a licence. `ollama` runs the model server
on host network (0.0.0.0:11434) and provides model-access at node scope, serving
its port and model — it mints nothing, so it is server-only, no runtime. The
`local-model-consumer` requires model-access, gets the endpoint (no secret), and
its templated openai.env carries OPENAI_BASE_URL=http://<at>:<port>/v1 +
OPENAI_MODEL. One provision, two answers: the mesh's own model behind the same
interface as a vendor's.

Proven end to end by the mesh-lab local-model bed (green).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 05:25:21 +02:00
jschoubben b33f9b14c7 Merge pull request 'openai-consumer: the static-key model-access consumer (ADR 0050)' (#18) from feat/openai-access into main 2026-09-07 04:25:56 +02:00
jschoubben c454c6a54c openai-consumer: the static-key model-access consumer (ADR 0050)
The consumer half of the OTHER model-access shape. Where anthropic-consumer
receives a refreshed access token, this receives one operator-supplied API key
the mesh sealed to it and the host unsealed at its secret path — no manager, no
refresh, no usage. It writes the key where an OpenAI/Codex client reads it: an
OPENAI_API_KEY env file and the publicly-known Codex auth.json. Pure node, no
SDK import — the simplest a model-access consumer gets.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 04:25:26 +02:00
jschoubben 4cc1b976f5 Merge pull request 'model-usage: the vendor-neutral usage store (ADR 0054)' (#17) from feat/usage-store into main 2026-09-07 04:08:21 +02:00
jschoubben e8c3ea8979 model-usage: the vendor-neutral usage store (ADR 0054)
The home ADR 0050 left open for a usage reading. mesh-control is a CLI and
cannot consume events, so the store that keeps the current usage picture is a
MODULE — the audit-logger's sibling: it consumes `module.*.usage.*` and upserts
each reading into its own provisioned postgres store, latest per
(licence, consumer, period, metric), in the clear. One vendor-neutral table
holds BOTH grains; they differ only in `consumer` (the holding module for the
licence grain, the session for the finer one). The consumer creates its table
on startup and, as a restart-until-ready service, self-heals rather than
gating the apply on a run-once that must reach a provider over the overlay.

The vendor->row normalisation moves into the adapter, as ADR 0054 requires:
anthropic-manager (licence grain, utilization%) and anthropic-consumer (session
grain, token/cost) now emit already-normalised { rows: UsageRow[], raw } on
their existing keys, so the store stays vendor-blind.

Proven end to end by the mesh-lab model-usage bed (green): a usage event
emitted into the mesh is upserted at both grains, latest-per-key, in the clear.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 04:07:26 +02:00
jschoubben bfbdf90496 Merge pull request 'anthropic model-access: manager + consumer modules' (#16) from feat/anthropic-module into main 2026-09-07 02:48:51 +02:00
jschoubben 19666ff054 anthropic-manager: seal over audited tweetnacl-sealedbox-js, not a hand-transcribed NaCl
The manager reseals a rotated refresh token to the node key with crypto_box_seal. That
seal was a full inline transcription of TweetNaCl's XSalsa20-Poly1305 and blakejs' BLAKE2b
(dependency-free, ~440 lines). Replace the internals with the audited
tweetnacl-sealedbox-js library — the same crypto_box_seal, on the same tweetnacl and blakejs
the mesh used to validate the seal during Phase C.

The exported API is unchanged: seal(value, recipientPublicB64) -> base64. The wire format is
unchanged too — ephemeralPub(32) followed by the box, nonce = blake2b(ephemeralPub +
recipientPub, 24) — so the host's Go box.OpenAnonymous still opens it. The mesh-control
cross-check fixture is regenerated from this seal().

The library and tweetnacl are added to the module's package.json dependencies so the runtime
image bundles them (blakejs arrives transitively). A local ambient .d.ts types the untyped
CJS bundle; it is imported as a default import because Node's ESM loader cannot see a UMD
bundle's named exports.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 02:17:37 +02:00
jschoubben 4c98bee043 anthropic-manager: seal the refresh token to the node key, do no crypto to open
The manager module drops its bespoke ECIES at-rest envelope and the node-private-key mount.
A module is never given a node's private key, so it cannot open an envelope -- the refresh
token is now delivered to it as cleartext by the host, unsealed from an ordinary sealed box.

  - sealedbox.ts: a dependency-free NaCl crypto_box_seal (node:crypto for X25519, transcribed
    XSalsa20-Poly1305 and BLAKE2b-24), byte-compatible with Go's box.SealAnonymous. It SEALS
    only -- opening is the host's job. Proven by a cross-language test in mesh-control.
  - adopt: reads the node's PUBLIC key from the delivered bound facts and seals the operator's
    refresh token to it, handing out only the box.
  - refresh: reads the refresh token as cleartext the host mounted, calls the vendor, re-seals
    a rotated token to the node's public key, submits only { access token, box }.
  - module.json: a model-access holder now -- binds the facts, binds the refresh token as a
    sealed secret; no keys dir, no MESH_NODE_SEALING_* mount.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 01:55:19 +02:00
jschoubben c206e2e11e anthropic model-access modules: manager (refreshable-grant) and consumer
Phase C of vendor-agnostic model-access (ADR 0050/0054). Two TypeScript
runtime modules:

- anthropic-manager: the refresh token is sealed at rest to the manager
  node's own key (atrest.ts, envelope encryption over X25519) and opened
  ONLY on the manager node. adopt seals the first envelope; refresh opens
  it, calls the Anthropic OAuth token endpoint, re-seals a rotated refresh
  token, and hands the control plane only the access token plus the opaque
  envelope. Also polls licence-grain usage (ADR 0054).
- anthropic-consumer: writes the delivered access token to
  ~/.claude/.credentials.json, access-token-only, atomically (the refresh
  token is never delivered); reports session-grain usage from the CLI
  transcripts; a fail-closed identity guard (expected-uuid plumbing is a
  flagged TODO).

Both run as scheduled containers (ADR 0053). Pure logic covered by
node --test fixtures (at-rest round-trip, credential strip, transcript
sum, refresh merge).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 01:00:12 +02:00
jschoubben cd46d53464 Merge pull request 'lavinmq: a user-facing AMQP provider (vhost-per-login)' (#15) from feat/convert-lavinmq into main 2026-09-06 23:21:27 +02:00
jschoubben 97707480d6 Add lavinmq amqp provider and amqp-ping consumer
lavinmq becomes a provider of a user-facing amqp interface: a consumer
that requires a message queue is given its OWN broker — a scoped vhost
and user on a lavinmq provider — not an account on the mesh's own
control-plane broker (ADR 0048). Vhost-per-login is the isolation model,
the exact analog of postgres's database-per-login: the provider names a
vhost after the consumer's login and a user with full rights on that
vhost and none elsewhere, so a login is a broker the consumer alone can
reach.

The provider drives lavinmq through its HTTP management API (client.ts,
the module's one impure seam), with a run-once bootstrap that computes
the RabbitMQ-compatible password hash lavinmq's config wants from the
plain admin secret the mesh mints — the value no ${secret:...}
placeholder can produce and the reason the bootstrap exists (ADR 0052).
serves.amqp carries the port so consumers reference ${bound:amqp:port}.

amqp-ping is a demo consumer: it contributes nothing (the vhost is the
login), reads its grant from an env-file the mesh fills, and uses
${bound:amqp:as} for BOTH its username and its vhost — the db-name
lesson applied to AMQP. It speaks AMQP 0-9-1 over a raw socket with no
npm dependency (the way redis speaks RESP) and round-trips one message.
It carries a slug so its identity fits the 20-char backend bound
(ADR 0049).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 15:50:19 +02:00
jschoubben 090a839d42 Merge pull request 'route-proxy: the native ingress module (drop traefik)' (#14) from feat/route-proxy-module into main 2026-09-06 15:17:21 +02:00
jschoubben 89254ef2c8 hello-web: add slug so its identity fits the 20-char backend bound (ADR 0049)
mesh_anchor_hello_web is 21 chars, one over the S3 access-key bound; slug 'hello' brings it to 17.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 15:16:56 +02:00
jschoubben 36640f68e6 Add route-proxy provider and hello-web consumer modules
route-proxy is the shipping form of the reference reverse proxy (novox/hq
ADR 0007, 08-connectivity section 3): it provides route, is given every
consumer as the file at receives.route, and forwards by the Host header. It
ships the Go proxy from mesh-control/examples/route-proxy via a multi-stage
Dockerfile; no broker, own-secret or provisioner, since it only reads the file
the mesh writes.

ACME_DIRECTORY defaults to Let's Encrypt staging and is overridable per node to
production, so there is no hardcoded production default -- resolving novox/hq
04-ISSUES/004. hello-web is a minimal consumer that requires route and
contributes name+port, to exercise the grant.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 15:07:16 +02:00
jschoubben 9dbbfbb5b4 Merge pull request 'fix: postgres consumers connect to the mesh-named db (the login), not a hardcoded name' (#13) from fix/consumer-db-name into main 2026-09-06 13:47:34 +02:00
jschoubben d301e05794 fix: postgres consumers connect to the db the mesh named (the login), not a hardcoded app name
The postgres provisioner creates each consumer a database named after the login the mesh
minted (mesh_<node>_<module>), per ADR 0048 -- 'a same-named database under exactly that
login'. But baserow, letta, umami, gitea, keycloak and nextcloud each hardcoded their app db
name (DATABASE_NAME=baserow, /letta, /umami, NAME=gitea, /keycloak, POSTGRES_DB=nextcloud),
so the service connected to a database that does not exist ('database letta does not exist').
Each now uses ${bound:postgres-database:as} as the db name -- the login, which is also the db
name -- matching the working meshboard pattern and the provisioner's actual behaviour.

Found by the two-node DB-consumer lab install (mesh-lab assigned-two-node-db); baserow is
proven connecting and running there. The S3 bucket name is the same class of assumption and is
a separate follow-up (s3 identity has its own length bound).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 13:47:16 +02:00
jschoubben dc12df18b4 Merge pull request 'fix: DB/cache/store providers must serve their port (migration blocker)' (#12) from fix/db-provider-serves-port into main 2026-09-06 03:31:50 +02:00
jschoubben 1f0c3ac69e fix: db/cache/store providers must serve their port; drop baserow's empty redis contribution
The breadth install (catalogue-broad) surfaced two latent resolution bugs that
single-module compile checks never caught (compile != resolve):

1. postgres/redis/mongodb/minio served no "port", yet seven consumers
   (baserow, letta, invoicing, gitea, umami, keycloak, nextcloud) reference
   ${bound:<provision>:port}. A binding auto-carries at/from/as; the port is the
   provider's half of the answer and must be declared in `serves` (manifest.go:
   "Serves is what a consumer needs to know ... a port, a path, a realm"). Added
   port to each: postgres 5432, redis 6379, mongodb 27017, minio 9000. Without it
   no DB consumer could resolve, let alone deploy.

2. baserow declared an empty contribution `contributes: {"redis-cache": {}}`.
   redis's serves spec for the cache is empty (a cache takes no per-consumer
   payload), so a consumer only `requires` it; an empty contribution is refused.
   Dropped it (requires/binds/secrets unchanged).

Found by mesh-lab catalogue-broad; the same fixes are mirrored in that bed.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 02:38:22 +02:00
jschoubben dba4554002 Merge pull request 'Convert confluence + jira on the tools-only template' (#11) from feat/convert-atlassian into main 2026-09-06 01:54:51 +02:00
jschoubben db8b60f4e4 Convert confluence and jira into nox catalog modules
Port the HAL confluence and jira integrations to the tools-only,
outbound-only external-SaaS pattern proven by the merged gitlab module:
runtime-only containers (container-runtime capability), own-secret token +
broker, a settings-managed config.json for public config, and a
per-module runtime image.

Each module carries its own Atlassian API client and tools (ADR 0039),
translated from HAL's @hal/sdk zod-schema/MCP-content shape into
mesh-sdk's input/run-returns-data shape. Public config (ATLASSIAN_URL,
ATLASSIAN_EMAIL) lives in config.json; the API token is the one
own-secret. Clients are built lazily and never throw at registration, so
each runtime serves its full tool surface with no credentials (the
Servarr lesson) — confluence serves 3 tools, jira serves 8.

jira's periodic ticket-poller (update-tickets.service/.timer) is NOT
ported: the mesh has no scheduled-task primitive yet (a pending
decision). Only jira's tools are ported; a top-of-file note records the
deferral.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 01:49:22 +02:00
jschoubben 2c3f241010 Merge pull request 'Convert gitlab — the tools-only external-SaaS exemplar (23 tools, lab-proven)' (#10) from feat/convert-gitlab into main 2026-09-06 01:40:31 +02:00
jschoubben 7eb156a82a Convert gitlab into a tools-only mesh catalog module
Port the HAL gitlab module (whose tools lived in @hal/sdk) into a
self-contained mesh-catalog module modelled on cloudflare-dns: the GitLab
API client and all its tools live in the module (ADR 0039), served through
mesh-sdk's registerModuleTools harness.

Tools-only, outbound-only external-SaaS shape: a runtime-only container on
network:host, no service, no listener, no provisioner. The token is an
own-secret; GITLAB_URL is a public setting in a merge:json config file.

The client is built lazily and never throws at registration, so the runtime
comes up and serves all 23 tools even with no valid token (the Servarr
lesson) — it only fails when a tool is actually invoked unconfigured.

Ported 23 tools: projects (list, get), merge requests (list, get, create,
approve, add note), pipelines (list, get, retry, cancel, list jobs, job log),
and project + group CI/CD variables (list, get, create, update, delete each).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 01:33:43 +02:00
jschoubben 09b140fe86 Merge pull request 'Convert baserow, letta, invoicing (DB/S3 consumers)' (#9) from feat/land-db-consumers into main 2026-09-06 01:14:30 +02:00
jschoubben 8b82e47889 invoicing: point app/api at the real registry host
The conversion left the image host as a placeholder (registry-api.example).
Point both containers at the actual registry the built app/api images live in
(novox/invoicing-{app,api}, tags 51-55 + latest present). Tracks :latest to
match the app's continuous-deploy model; pin by digest later if reproducibility
of a specific build is wanted.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 01:14:08 +02:00
jschoubben 515b1e79c7 Convert baserow, letta, invoicing to mesh catalog DB consumers
Mirror the postgres-consumer shape (keycloak/nextcloud): requires the
backing service(s), binds/secrets for the delivered credential, and a
server container that reads it from an interpolated env file.

- baserow: consumes postgres-database + redis-cache; tooled (dormant
  until an account is configured, like gitea's token).
- letta: consumes postgres-database; tooled, live via a mesh-minted
  server-password injected into both server and runtime.
- invoicing: consumes mongodb-database + s3-bucket; plain two-container
  service (app + api), no tools.

Digests pinned for baserow and letta; invoicing keeps private-registry
tags (DIGEST-UNRESOLVED). redis-cache and mongodb-database consumer
shapes are inferred (no prior consumer in the catalog).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 01:12:50 +02:00
jschoubben 3b5f214f3d Merge pull request 'mosquitto: run-once dynsec bootstrap + fix the exit-0-on-failure provisioner' (#8) from feat/mosquitto-bootstrap into main 2026-09-06 01:06:08 +02:00
jschoubben e58cbc1c46 mosquitto: don't read a rejected dynsec command as success
mosquitto_ctrl's dynsec subcommands exit 0 even when they fail — a
"Client not found", an "already exists", a "Connection error: Not
authorized", an "Unable to connect" all return status 0 and report the
failure only as a line of text (verified live against 2.0.11). ctl()
trusted the exit code, so clientExists()'s getClient probe never threw
and always returned true; createScopedClient therefore took the
setClientPassword branch, never ran createClient, and the consumer's
client never landed in the dynsec store — while the provisioner logged
it as provisioned. That is the assigned-catalogue-mqtt failure.

ctl() now scans the combined stdout/stderr for the tool's error markers
and raises a match as the failure it is. Surfacing those errors exposed
two calls that only "worked" by being swallowed: addRoleACL re-run
reports "already exists" (now ignored like createRole), and addClientRole
re-run reports a bare "Internal error" that cannot be told from a real
fault — so the binding is checked with clientHasRole and only added when
absent. Fresh provision, idempotent re-provision, password rotation, bad
admin auth and unreachable broker all verified against a live broker.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 00:57:45 +02:00
jschoubben ec1e7189a3 mosquitto: seed dynsec with a run-once bootstrap before the broker
The Dynamic Security plugin refuses to start the broker unless
dynamic-security.json already holds an admin client, and no reconcile loop
seeds it (novox/hq ADR 0052). Add a run-once init container, declared before
the server container, that runs mosquitto's own bootstrap entrypoint in the
runtime image: it seeds the store offline via mosquitto_ctrl and exits, and
the host gates the broker on its completion.

The bootstrap hands the seeded file to the broker's user (uid 1883, chown +
0600): the broker must read the seed at startup AND persist to it as clients
come and go, but the init container runs as root and would otherwise leave a
file the broker can neither read nor rewrite. This is the ownership question
ADR 0052 left for the lab to settle. It seeds only when the file is absent, so
what the running plugin grows is never clobbered (issue 035).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 00:33:28 +02:00
jschoubben eb5cf65be3 Merge pull request 'Servarr events entrypoints don't crash the runtime when unconfigured' (#7) from fix/arr-events-guard into main 2026-09-05 23:14:43 +02:00
jschoubben 419dd68145 Servarr events: don't crash the runtime when unconfigured
The radarr/sonarr/lidarr/bookshelf events entrypoints built their client
with XClient.fromEnv() at import time, which throws when no API key/URL is
available yet — crash-looping the runtime container. The tools already
guard this; the events entrypoint did not.

Mirror the tools' guard: build the client in a try/catch, and only start
the poll loop when it succeeds. When it fails, log one line and stay idle
until a key is available. Behaviour when configured is unchanged.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 23:13:56 +02:00
jschoubben 50eb1e6f02 Merge pull request 'Media modules access the shared library, they do not own it (ADR 0051)' (#6) from feat/media-access-not-ownership into main 2026-09-05 22:48:16 +02:00
jschoubben 69c8a5edf8 media stack: the shared library is accessed, not owned
04-ISSUES/036: eight media modules each declared the shared library and
download directories under /services/media/* as their own `directory`
resources. Six of them owning one path is the collision the resolver
refuses — so the stack's only sensible assignment, all on one machine
sharing one filesystem, would be refused the first time two landed
together.

The media library is the operator's, owned by no module (novox/hq
ADR 0051). Move every /services/media/* path from an owned `directory`
resource to an `accesses` entry: the mesh mounts it and owns nothing —
does not create, chown, reconcile or remove it — and several modules may
access one path with no conflict. Each module's own config and mesh-state
directories stay owned resources.

Modes are least-privilege: plex reads the libraries it streams; the
managers and download clients get read-write on what they import and
write; bazarr writes subtitles into the libraries (read-write) and only
reads the download spool. The container volume mounts are unchanged.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 22:14:33 +02:00