Commit Graph
191 Commits
Author SHA1 Message Date
jschoubben 9387f8b040 A module that listens is served, not run
`run` imports an entrypoint without binding a broker — it exists for a step that
works offline and exits. Both the catalogue and amqp-ping subscribe on import,
so both died on the first on() with no broker bound.
2026-09-13 01:30:39 +02:00
jschoubben c774d5dbe0 Install the postgres client the way the runtime base can
The published base is debian; apk is not there and the build said so.
2026-09-13 01:28:15 +02:00
jschoubben 6ebf51d312 postgres's runtime carries the client it provisions through
Its provisioner runs DDL by shelling out to psql, which the runtime base has no
reason to hold. Every create failed with ENOENT and retried for ever.
2026-09-13 01:26:23 +02:00
jschoubben 594295f009 The catalogue restarts when its database credentials change
Without it the container keeps whatever the env file said when it was created.
Nothing reports that: it runs, and it is wrong.
2026-09-13 01:22:38 +02:00
jschoubben f7d57e9556 The catalogue brings its own postgres driver
The runtime base carries what every module needs, and a database driver is not
that. Installed into an empty directory because the module's package.json also
names the sdk, which lives in the base rather than on a registry.
2026-09-13 01:16:23 +02:00
jschoubben e742b6a569 The catalogue builds its own runtime, and postgres serves its tools
Both modules keep their tools in an entrypoint of their own, so an image that
named only the consumer would serve none of them.
2026-09-13 01:13:43 +02:00
jschoubben d6c9c8d666 postgres builds its own runtime, like any other module
Its provisioner container named an image nobody could produce — a zero digest
placeholder. It names an artifact instead, and the module says how to build it,
so the mesh can make the database provider the catalogue needs.
2026-09-13 01:10:45 +02:00
jschoubben 2a6fed6f4f The builder is told the mesh's name for the machine it runs on 2026-09-13 01:07:37 +02:00
jschoubben 81a80c675c The builder declares what it announces
Its account is scoped from what it emits and consumes, and it declared neither —
which is why asking for a generic module account produced one that authenticated
and could do nothing, with the refusal surfacing a layer away as a permissions
error against a queue.

Declaring the announcement is not documentation here. It is what the permission
is derived from.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-13 00:57:41 +02:00
jschoubben 39631d6f87 amqp-ping says what it is made of, and can be built from its own directory
A module's runtime image was assembled by a script copying the sdk and the tool
runtime out of neighbouring checkouts, so it could only be built on a workstation
that had them. That is why no module declared what it was made of and why
forty-seven point at a placeholder.

The tool runtime becomes an image a module's runtime is built FROM, published like
any other artifact. The module then builds from its own directory and that base —
one clone, which is what the builder can actually be asked for (novox/hq ADR 0069).
The dependency stops being a property of somebody's machine and becomes a build
edge, pinned to a digest the mesh's registry assigned.

The compiler is invoked by its real path rather than through node_modules/.bin:
those are symlinks to a launcher that requires its library relatively, and
resolving them while building the base leaves a launcher pointing at nothing.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-13 00:52:57 +02:00
jschoubben a73cb8a2f8 The catalogue, as a module that owns the module graph
It links module-versions to each other and knows nothing about nodes; which
machine runs what stays the control plane's (novox/hq ADR 0070, 0072). Keeping
them apart is what lets the control plane carry on composing declarations while
this is down.

The builder announces what it built, this places it in the graph and announces
what that means, and the control plane hooks the meaning rather than the build
output. A rebuild producing the commit already current is registered and is not
an upgrade — announcing it would ripple outward forever through modules that did
not change.

Ordering is not computed. Modules stale and waiting on nothing that is itself
stale are announced as buildable; the rest stay stale and appear once whatever
they were waiting for is registered, so a chain and a diamond need no special
handling and nothing holds a plan.

Four tools over the graph: what this mesh holds, one module in full, what a
change to a module reaches, and what must be rebuilt and why. The edges are
derived from builds rather than declared, so they cannot drift from what the code
actually uses.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 23:16:05 +02:00
jschoubben a4d10341e7 Declare what hello-web is made of, to prove a build from a path
A proof branch, not for main: the route-forwarding bed reads this module's literal
image and would break until the module is built.

The modelling is right regardless — the image is upstream, so the mesh should
mirror it once into its own registry and pin what that registry assigned, rather
than every machine fetching a reference somebody else can move.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:57:18 +02:00
jschoubben 5f76e34994 Add the builder as a module, so the mesh can be given one
The code that turns a repository into artifacts already existed and is deliberately
something a machine runs as an ordinary module rather than something the control
plane does. There was no module for it, so it could never be placed and never ran —
which is why nothing in the catalogue could be produced.

It asks for a container runtime because it builds images, requires the artifact
store because it publishes into it, and claims one per machine. Its credential
folder is mounted rather than the file, since a file mount keeps pointing at the
old contents after the mesh writes new ones.

The store is reached on the machine's own loopback: the address in a binding is
where a machine sits on the private network, and the store has to be co-located
anyway because runtimes refuse a plain-HTTP registry anywhere else.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:45:59 +02:00
jschoubben 5aacd2538b mesh-control: the enrolment endpoint is the substrate's, not the overlay's
The module said MESH_BROKER_ADDRESS=${machine:at}:5671, and that cannot work in either
direction.

At genesis it does not resolve at all. `at` is a machine's name on the PRIVATE network,
and the mesh only holds one for a node that has an overlay placement and resolves the
networking module — neither of which exists when the control plane is installed, which
is step 9 of ten, long before anything has been placed anywhere. mesh-control refuses a
${machine:} key it does not hold rather than writing the literal through, so the push
would have stopped with "this machine says name".

And afterwards it would be the wrong address anyway. This value is what every enrolment
token tells a joining node to dial. A machine that has not enrolled is not on the
overlay, so an overlay name is precisely the one thing it cannot reach.

It is the substrate's own fact — the address the broker advertises, decided by whoever
wrote the bundle, which the mesh did not make and cannot invent. So it arrives the way
the store connections beside it arrive: an own-secret the installer delivers with
`secret accept`, read out of the bundle it produced. mesh-bootstrap already does this for
every variable the module fills from a secret; this one simply joins them.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 11:48:51 +02:00
jschoubben 2c322cb2fb mesh-control: the control plane could not read its own connections
Its image is FROM scratch and runs as 65534. The host writes a sealed own-secret 0600,
owned by root, which is right — but the module then bind-mounted those three files into
the container and told the process to open them. It cannot:

  $ docker run --rm -v <0600 root file>:/run/secrets/inventory:ro \
      -e MESH_STORE_INVENTORY_FILE=/run/secrets/inventory mesh-control:development status
  MESH_STORE_INVENTORY_FILE names /run/secrets/inventory ... and it cannot be read:
  open /run/secrets/inventory: permission denied

Measured on a workstation, not reasoned about. Every other module in this catalogue gets
away with the same mount because its runtime container runs as root; this one does not,
and genesis (novox/hq ADR 0067) would have stopped at step 9 with a control-plane module
that starts and cannot open a context.

The connections go through the env file this module already has instead. That file is
mode 0600 and is read by the container runtime's client, which is root — the same reason
the broker's URL has always reached the process this way. It also sidesteps the inode
that a file bind mount pins (290be37, step-ca): --env-file is read afresh at create, and
restart-on names it.

The own-secrets stay exactly as they were, because the installer delivers the substrate's
real connection strings into them with `secret accept` before the first push — the mesh
did not make those credentials and cannot invent them.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 11:38:28 +02:00
jschoubben 290be37a93 step-ca: mount the directory, not each secret file
A file bind mount tracks the inode. The host writes atomically — new file, rename
over — so the container keeps reading the file that was there when it started,
and a rotated or newly-delivered secret never reaches it. Mounting the parent
directory resolves the path on each open instead.

This is a known shape (hal KB troubleshooting/docker-bind-mounts), and it cost an
hour here before it was looked up: the CA crash-looped on a root key it had
already been given, because the container still held the inode from before the
key arrived.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 08:49:55 +02:00
jschoubben 0d2c2dd989 minio: pin the manifest these machines can actually run
The pinned digest named the arm64 manifest, so an x86 machine pulled it and the
container died with 'exec format error' on every restart. It never showed while
the lab ran its own registry: the harness pushed the WORKSTATION's copy, which
is amd64, and every machine then pulled that under a digest the registry had
just assigned. The registry was quietly correcting the architecture too.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 02:51:34 +02:00
jschoubben 9ff320ecca Pin the operator's own images by digest, as every other image already is
Nine references across seven modules named ':latest'. ADR 0006 forbids it and
the host refuses it by name — and the refusal had never fired, because the lab
pushed every image into its own registry and rewrote each reference to the
digest it had just assigned. Deleting that registry made these the only
manifests the host would now reject (novox/hq 04-ISSUES/039).

The digests are what each tag resolves to today, read from the registry that
serves them.

This is a stopgap and should be said as one: a digest written into a repository
is wrong the moment anybody rebuilds, which is precisely why the design has the
repository name artifacts and the mesh hold digests. Until something builds and
publishes, a digest that is stale is still better than a tag that silently moves.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 01:06:14 +02:00
jschoubben 4a1c5ac817 The artifact store requires nothing, because at genesis there is nothing
A route-label migration gave the registry a public name, and with it a
requirement. But the registry is the first module a new mesh installs: at that
moment nothing provides a route, so the push is refused and a mesh cannot get
its own image store — the cycle 04-ISSUES/029 closed, re-entered through a
different door.

The same rule that record states applies: a module providing the artifact store
cannot depend on what the store is needed to deliver. A public name for it is an
ordinary want and belongs to a module beside it, installed once there is a mesh
to install things.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:06:23 +02:00
jschoubben 71f8012c5b The control plane as an ordinary module
novox/hq ADR 0067 pivots genesis through a temporary control plane and then
reinstalls the control plane as an ordinary module pinned to a digest the mesh's
own registry assigned. That record notes the one thing missing: a control-plane
module manifest, which did not exist.

It could not be written honestly before now. The control plane read its store
connection from MESH_STORE_<CONTEXT>, that connection string carries a password,
and a manifest can put a sealed value into a file's `content` but has nothing
that substitutes into a container's `env`. So the manifest could carry the
password in the clear, or omit the setting. mesh-control now also accepts
MESH_STORE_<CONTEXT>_FILE, which is how every other module here is given secret
material, and the manifest follows.

What the substrate bundle gives the control-plane container today, and where
each part has gone:

  MESH_STORE_INVENTORY   own-secret `inventory`, mounted, named by _FILE
  MESH_STORE_IDENTITY    own-secret `identity`,  mounted, named by _FILE
  MESH_STORE_LICENCES    own-secret `licences`,  mounted, named by _FILE
  MESH_BROKER_AMQP       own-secret `broker`, through an env-file hole
  MESH_BROKER_MANAGEMENT own-secret `broker-management`, likewise
  MESH_BROKER_ADDRESS    ${machine:at}:5671 in that same env-file
  MESH_BROKER_CERTIFICATE  plain env; the path is not a secret
  network host, args ["serve"], the broker's TLS volume  unchanged

The two broker URLs go through an env-file rather than a file of their own
because mesh-control has no MESH_BROKER_AMQP_FILE. That is the same fault one
layer over, and the same remedy would fix it; it is out of this change's scope
and is written down rather than papered over.

None of these values is in the manifest. Each is an own-secret the operator
supplies with `secret accept` — the mesh cannot invent a connection string — and
the container restarts when any of them changes.

The module claims `the-control-plane` at mesh scope, which the bundle has no way
to say: two control planes writing one inventory is a fault worth refusing at
assignment. It carries no `listens`, because `serve` dials the broker and binds
nothing. The image is the catalogue's placeholder digest for a mesh-built image,
which the installer replaces with what the registry assigned.

Checked with the real parser: all 67 manifests through catalogue.ParseManifest
and every module's CheckIdentity against all four node names — 0 problems — and
this manifest rendered through Resolution.Declaration, so the ${secret:…} names,
${machine:at}, the restart-on ids and the image pin are exercised rather than
merely parsed. Slug `control`: mesh_shanks_control is 19 of the 20 an S3 access
key keeps.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:43:07 +02:00
jschoubben 16b4dc9ab9 step-ca: certify the machine the mesh reaches it at
Its API certificate carried localhost only, so a proxy dialling the address the
mesh handed over refused it on hostname verification. The names now compose from
the machine the module was assigned to, which a manifest could not know and now
does not have to (mesh-control ${machine:...}).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:12:48 +02:00
jschoubben d9fbc72565 step-ca: give the CA's own secret material to the uid the CA runs as
The upstream `smallstep/step-ca` image runs as uid 1000. An own-secret lands as
a file the host writes root:root 0600 — the module says only a name and a path,
so there is nowhere to say who must be able to read it — and the container
crash-looped on `permission denied` reading its own root key. The lab got past
it first with `chmod 0644`, which hands the root key to every local user, and
then by running the CA as root, which is worse.

Neither is needed. A module composes its own files, and a `file` resource takes
both a `mode` and an `owner`, with `${secret:name}` reaching the module's own
secrets — the mechanism redis already uses to hand its password to a server
running as 999. So the three pieces of init material are declared as owned
files: still 0600, owned by 1000:1000, and those are what the container mounts.

The raw own-secret files stay where they were. Nothing mounts them now; they are
how the secret comes to exist on the machine, and the host is the only thing
that reads them. Same shape as redis's `default.secret`.

Verified against the real code rather than by inspection: mesh-control attaches
the sealed value to each of the three resources and keeps `owner` and `mode`
(`sealedFor` + `intoFile` over this manifest), and mesh-host parses `owner` on a
sealed-substituted file and lands it 0600 owned by 1000:1000 (`applyFile`).
2026-09-10 20:57:58 +02:00
jschoubben 43ca9c9c73 Give every module that would overflow its login a short slug
An identity is `mesh_<node>_<slug-or-name>` and a backend keeps 20 characters
(an S3 access key). Overflow makes a module unresolvable, and this catalogue
was finding it one module at a time, on a raise: route-proxy on novox is 22,
home-assistant on ace is 23. Two found by hand where a sweep would have found
eighteen.

So the whole catalogue was swept instead, against the longest node name the
mesh actually has (`shanks`, six characters) rather than against the node each
module happens to sit on today — a module is assigned somewhere, and where is
not a property of the manifest. That leaves eight characters for the identity
source, and eighteen modules were over it.

Slugs added, chosen to stay greppable in a provider's user list:

  anthropic-consumer  claude     openai-consumer     openai
  anthropic-manager   anthmgr    portainer           portain
  audit-logger        audit      public-acme         pubacme
  bookshelf           books      qbittorrent         qbt
  cloudflare-dns      cfdns      resolv-conf         resolv
  confluence          confl      resolved-split-dns  splitdns
  home-assistant      hass       route-proxy         rproxy
  invoicing           invoice    verdaccio           verdacc
  mosquitto           mosq
  nextcloud           ncloud

A slug changes the login the mesh mints, so a module already provisioned under
its full name is re-minted under the slug and its old login withdrawn — which
is the provisioner's ordinary business, but it is a change, not a no-op.

Checked with the real parser: every one of the 66 manifests through
`catalogue.ParseManifest`, and every module's `CheckIdentity` against all four
node names. 0 problems, where the same check over the parent commit reports 44.
2026-09-10 20:57:45 +02:00
jschoubben 08be6683b6 route labels: nextcloud is 'drive' (real hostname), novox.be is the apex '@'
nextcloud's label was migrated from a wrong module-name default; its real
production hostname is drive.novox.be. novox.be is the bare-domain apex, now the
'@' label (composeName gained apex support).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 00:27:25 +02:00
jschoubben f0665ba956 ADR 0056: selectable ACME issuer, step-ca init from operator root, route labels
Piece A + B of ADR 0056, completing the internal-CA work in ff01ada.

Selectable issuer. acme-ca is now a role two providers can satisfy: step-ca
(internal CA) or the new public-acme (a fact-only module, no container/listen)
that serves Let's Encrypt production. A mesh assigns one or the other to satisfy
route-proxy's `requires: acme-ca`.

One directory shape for both. A provider serves the ACME directory's parts the
way the mesh already models any reachable service -- an address (`at`), a `port`
and a `path` -- and route-proxy composes `https://<at>:<port><path>`. step-ca
lets the mesh fill `at` (its node) and `port` (its single listen) and serves only
`path`; public-acme, not being a mesh service, serves all three (overriding `at`
with the public host). Same composition either way.

Empty root means the system trust store. Both providers serve `root`: step-ca
the operator root PEM (settled per mesh), public-acme an empty string. route-proxy
writes it to the CA bundle file unconditionally; the binary now reads an empty
bundle as "the root is already trusted by the OS" and falls back to system roots
(examples/route-proxy/main.go, committed on the mesh-control ADR-0056 branch).

step-ca inits from the operator's root. The operator's root cert, root key and
root-key password are mounted at the smallstep entrypoint's default init paths
(/run/secrets/root_ca.crt, root_ca_key, root_ca_key_password) with the matching
DOCKER_STEPCA_INIT_*_FILE vars, so `step ca init` adopts the operator's root
instead of self-generating one -- the CA that signs is the CA route-proxy trusts.

Route names are labels, not FQDNs. Every routed module now contributes a `label`
(the leftmost subdomain) instead of a full public hostname; the node's public
domain composes the name. Apex (novox.be) is left as a full name -- composeName
has no empty-label/apex convention yet (mesh-control follow-up).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 00:21:42 +02:00
jschoubben ff01adaa45 Add internal ACME CA (step-ca) as a generic acme-ca provision; wire route-proxy to consume it
Piece C of ADR 0056: an internal authority certifies routed names by the same
path a public one would, with the proxy pointed at whichever issuer the mesh
names and trusting that issuer's root.

step-ca (new): provides acme-ca (mesh scope), listens 9000 from mesh, persists
its CA home under uid 1000. The root CA cert reaches consumers as a served
value settled from a per-mesh operator setting (no baked root); the init
password and root key are sealed secrets, not literals. No route-name -> IP
hosts map (that is Piece B). Upstream smallstep/step-ca pinned by Docker Hub
digest.

route-proxy: requires + binds acme-ca. ACME_DIRECTORY is composed on the
consumer side from ${bound:acme-ca:at}:${bound:acme-ca:port}, and ACME_CA_BUNDLE
is a file whose content is ${bound:acme-ca:root} -- both from the binding,
replacing the hardcoded directory and the baked root PEM.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 23:07:54 +02:00
jschoubben d018f1e16e Route names mirror real production hostnames from live Traefik
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 14:04:17 +02:00
jschoubben e26ca38eaa novox conversions: give the web apps public names through route-proxy
Add a `route` contribution (requires/contributes/binds) to every web app so
each gets a Host-routed public name via route-proxy, mirroring the
de-spiegel/only-office pattern:

- novox: gitea, keycloak, nextcloud, umami, invoicing, verdaccio, registry,
  and mailu (single mail.novox.be -> 7080; admin/webmail/api ride that port).
- ace: grafana, sonarr, radarr, lidarr, bazarr, ombi, tautulli, jackett,
  nodered, searxng, home-assistant, bookshelf, baserow.

Split photos so its three sites each get a name: photos keeps server +
admin-client (photos.novox.be), and new photos-eef (eef.novox.be) and
photos-filip (filip.novox.be) modules carry the client sites.

The production FQDN stays the literal default; a per-node .incus name is a
settings override applied where the mesh runs, not a manifest hardcoding.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 12:51:15 +02:00
jschoubben 431310fb03 novox conversions: short slugs for the login-length cap + real roundcube digest
The whole-mesh-novox lab validation found three modules unresolvable: a
provision-consuming module's minted login is mesh_<node>_<module>, and on novox
`mesh_novox_only_office`(22)/`_de_spiegel`(21)/`_amqp_email_forwarder`(31)
overflow the 20-char S3 access-key cap (ADR 0049). Added a short `slug` each
(office/spiegel/emailfwd → 17/18/19 chars). Also fixed mailu's roundcube image:
the pinned digest returns "manifest unknown" from ghcr; corrected to the real
:1.9 digest.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 02:35:47 +02:00
jschoubben b3309a0ce9 mailu: full multi-container stack from the live novox deployment
Rebuild the mailu manifest from the running production deployment as the
source of truth, and finish wiring its tool runtime.

- Full 11-container topology: front, smtp, imap, admin, antispam, antivirus,
  webmail, webdav, fetchmail, resolver, redis — real ghcr.io/mailu images
  pinned by digest at tag 1.9 (clamav/radicale/fetchmail digests newly fetched).
- Admin DB now consumes the mesh postgres-database provider (requires +
  contributes + binds + secrets, DB_* templated from ${bound}/${secret}),
  replacing the bundled postgres:13 admindb the live stack still runs.
- Full config env from the live containers as a plain env-file; SECRET_KEY,
  the admin API token and the initial-admin password become own-secrets;
  DB password comes from the provider secret. No secret values hardcoded.
- Enable the Mailu admin REST API (API=true, WEB_API, API_TOKEN own-secret) so
  the ported tools can reach it — the live deployment runs this API OFF.
- mesh-mailu tool runtime on the mailu network: admin API over the module
  network, token from the mounted own-secret, docker.sock for the doveadm mail
  reads, broker + mergeable config with restart-on.
- client.ts fromEnv reads the API token from its mounted own-secret file
  (MESH_MAILU_API_KEY_FILE), matching the cloudflare-dns/umami pattern.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 01:04:55 +02:00
jschoubben 8f99b5bd42 Add amqp-email-forwarder daemon module
Owned daemon that consumes from the mesh amqp provision (lavinmq) and
forwards to SMTP. No ports, no route. AMQP host/port/user/password
templated from the amqp grant (analog: amqp-ping); vhost/exchange/queue
are the app's own config; SMTP user+password are given own-secrets
(secret accept). See report for the EMAILDELIVERY_T vhost integration flag.
2026-09-09 00:41:06 +02:00
jschoubben c423e21a8e Add novox.be; replace photos immich stub with the real app
- novox.be: owned www:latest container, public route novox.be
  (host 4000 -> app 8080). No DB, no secrets. Analog: de-spiegel/hello-web.
- photos: the existing module.json was a wrong immich stub (alpine image,
  port 2283). Replace it with the user's real photo app: photos-server
  backend consuming the mesh s3-bucket (bucket photos) + mongodb-database
  (db photos) providers, env templated from the provisions (analog:
  invoicing), plus the three static client containers (admin + two family
  sites). Only photos.novox.be is routed: contributes.route is single-valued
  across the catalog, so eef/filip need their own modules (flagged).
2026-09-09 00:36:48 +02:00
jschoubben 1fc10be11b Convert only-office, n8n, de-spiegel to nox modules
Mirror the proven converted modules: pure manifests, no runtime.

- only-office: owned onlyoffice/documentserver container, generated JWT
  own-secret, public route office.novox.be. Analog: gitea/hello-web.
- n8n: consumes mesh postgres-database + redis-cache providers instead of
  bundling them, owned container, public route n8n.novox.be. Analog:
  baserow (postgres+redis) + hello-web (route).
- de-spiegel: owned container built from its own repo image, public route
  de-spiegel.novox.be, given SMTP own-secrets. Analog: hello-web/invoicing.
2026-09-09 00:30:06 +02:00
jschoubben 5973d41966 umami: read the admin password from its mounted secret file
The whole-mesh dry-run found umami's runtime crash-looping "admin password is
not set": its `admin` own-secret is mounted at /run/secrets/admin, but the
client read the bare env UMAMI_ADMIN_PASSWORD, which nothing sets. Same shape as
the six tool-runtime credential fixes — read the mounted file first
(MESH_UMAMI_ADMIN_PASSWORD_FILE), falling back to the env. (photos and mailu
remain deeper conversion jobs — a stub app image and a full Mailu config env —
not credential-wiring, tracked separately.)

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-08 18:43:48 +02:00
jschoubben d289e5a928 modules: wire tool-runtime app credentials as own-secrets
The six modules that run a mesh-<mod> tool-runtime sidecar read an app
credential from an env var the manifest never provided, so the sidecar
crash-looped in the whole-mesh dry-run (e.g. "no Plex token — set
MESH_PLEX_TOKEN"). These are operator-set app secrets, so deliver them the
same way cloudflare-dns delivers its API token: an own-secret file mounted
read-only, with a MESH_<APP>_*_FILE env pointing at the mount, and the
runtime code preferring that file (falling back to the existing env so
nothing regresses).

- plex: own-secret token -> /run/secrets/token, MESH_PLEX_TOKEN_FILE
- bazarr: own-secret api-key -> /run/secrets/api-key, MESH_BAZARR_API_KEY_FILE
- ombi: own-secret api-key -> /run/secrets/api-key, MESH_OMBI_API_KEY_FILE
- home-assistant: own-secret token -> /run/secrets/token, MESH_HOMEASSISTANT_TOKEN_FILE
- nzbget: own-secret password -> /run/secrets/password, MESH_NZBGET_PASSWORD_FILE (URL stays plain env)
- qbittorrent: own-secret password -> /run/secrets/password, MESH_QBITTORRENT_PASSWORD_FILE (URL stays plain env)

The operator now completes each with `secret accept <node> <module> <name> --from <file>`.
tsc passes for all six.
2026-09-08 18:40:23 +02:00
jschoubben 75fb16bbfb fail2ban: require the firewall capability, not the non-existent intrusion-prevention
The whole-mesh dry-run found fail2ban unassignable on every node: it declared
`capabilities: ["intrusion-prevention"]`, which mesh-host has no detector for
(its detectors are container-runtime, package-manager, service-manager,
firewall, overlay, graphical-session, seat, privileged). intrusion-prevention
is what fail2ban PROVIDES, not a host capability it needs. It bans via
iptables/ufw, so it needs `firewall` — the same capability the firewall module
declares. The `the-intrusion-prevention` claim (node-exclusive) is unchanged.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-08 18:27:34 +02:00
jschoubben 29fe7c5bd0 local model: an ollama provider + a local-model consumer (ADR 0055)
Model access answered by a NODE, not a licence. `ollama` runs the model server
on host network (0.0.0.0:11434) and provides model-access at node scope, serving
its port and model — it mints nothing, so it is server-only, no runtime. The
`local-model-consumer` requires model-access, gets the endpoint (no secret), and
its templated openai.env carries OPENAI_BASE_URL=http://<at>:<port>/v1 +
OPENAI_MODEL. One provision, two answers: the mesh's own model behind the same
interface as a vendor's.

Proven end to end by the mesh-lab local-model bed (green).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 05:25:21 +02:00
jschoubben c454c6a54c openai-consumer: the static-key model-access consumer (ADR 0050)
The consumer half of the OTHER model-access shape. Where anthropic-consumer
receives a refreshed access token, this receives one operator-supplied API key
the mesh sealed to it and the host unsealed at its secret path — no manager, no
refresh, no usage. It writes the key where an OpenAI/Codex client reads it: an
OPENAI_API_KEY env file and the publicly-known Codex auth.json. Pure node, no
SDK import — the simplest a model-access consumer gets.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 04:25:26 +02:00
jschoubben e8c3ea8979 model-usage: the vendor-neutral usage store (ADR 0054)
The home ADR 0050 left open for a usage reading. mesh-control is a CLI and
cannot consume events, so the store that keeps the current usage picture is a
MODULE — the audit-logger's sibling: it consumes `module.*.usage.*` and upserts
each reading into its own provisioned postgres store, latest per
(licence, consumer, period, metric), in the clear. One vendor-neutral table
holds BOTH grains; they differ only in `consumer` (the holding module for the
licence grain, the session for the finer one). The consumer creates its table
on startup and, as a restart-until-ready service, self-heals rather than
gating the apply on a run-once that must reach a provider over the overlay.

The vendor->row normalisation moves into the adapter, as ADR 0054 requires:
anthropic-manager (licence grain, utilization%) and anthropic-consumer (session
grain, token/cost) now emit already-normalised { rows: UsageRow[], raw } on
their existing keys, so the store stays vendor-blind.

Proven end to end by the mesh-lab model-usage bed (green): a usage event
emitted into the mesh is upserted at both grains, latest-per-key, in the clear.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 04:07:26 +02:00
jschoubben 19666ff054 anthropic-manager: seal over audited tweetnacl-sealedbox-js, not a hand-transcribed NaCl
The manager reseals a rotated refresh token to the node key with crypto_box_seal. That
seal was a full inline transcription of TweetNaCl's XSalsa20-Poly1305 and blakejs' BLAKE2b
(dependency-free, ~440 lines). Replace the internals with the audited
tweetnacl-sealedbox-js library — the same crypto_box_seal, on the same tweetnacl and blakejs
the mesh used to validate the seal during Phase C.

The exported API is unchanged: seal(value, recipientPublicB64) -> base64. The wire format is
unchanged too — ephemeralPub(32) followed by the box, nonce = blake2b(ephemeralPub +
recipientPub, 24) — so the host's Go box.OpenAnonymous still opens it. The mesh-control
cross-check fixture is regenerated from this seal().

The library and tweetnacl are added to the module's package.json dependencies so the runtime
image bundles them (blakejs arrives transitively). A local ambient .d.ts types the untyped
CJS bundle; it is imported as a default import because Node's ESM loader cannot see a UMD
bundle's named exports.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 02:17:37 +02:00
jschoubben 4c98bee043 anthropic-manager: seal the refresh token to the node key, do no crypto to open
The manager module drops its bespoke ECIES at-rest envelope and the node-private-key mount.
A module is never given a node's private key, so it cannot open an envelope -- the refresh
token is now delivered to it as cleartext by the host, unsealed from an ordinary sealed box.

  - sealedbox.ts: a dependency-free NaCl crypto_box_seal (node:crypto for X25519, transcribed
    XSalsa20-Poly1305 and BLAKE2b-24), byte-compatible with Go's box.SealAnonymous. It SEALS
    only -- opening is the host's job. Proven by a cross-language test in mesh-control.
  - adopt: reads the node's PUBLIC key from the delivered bound facts and seals the operator's
    refresh token to it, handing out only the box.
  - refresh: reads the refresh token as cleartext the host mounted, calls the vendor, re-seals
    a rotated token to the node's public key, submits only { access token, box }.
  - module.json: a model-access holder now -- binds the facts, binds the refresh token as a
    sealed secret; no keys dir, no MESH_NODE_SEALING_* mount.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 01:55:19 +02:00
jschoubben c206e2e11e anthropic model-access modules: manager (refreshable-grant) and consumer
Phase C of vendor-agnostic model-access (ADR 0050/0054). Two TypeScript
runtime modules:

- anthropic-manager: the refresh token is sealed at rest to the manager
  node's own key (atrest.ts, envelope encryption over X25519) and opened
  ONLY on the manager node. adopt seals the first envelope; refresh opens
  it, calls the Anthropic OAuth token endpoint, re-seals a rotated refresh
  token, and hands the control plane only the access token plus the opaque
  envelope. Also polls licence-grain usage (ADR 0054).
- anthropic-consumer: writes the delivered access token to
  ~/.claude/.credentials.json, access-token-only, atomically (the refresh
  token is never delivered); reports session-grain usage from the CLI
  transcripts; a fail-closed identity guard (expected-uuid plumbing is a
  flagged TODO).

Both run as scheduled containers (ADR 0053). Pure logic covered by
node --test fixtures (at-rest round-trip, credential strip, transcript
sum, refresh merge).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 01:00:12 +02:00
jschoubben 97707480d6 Add lavinmq amqp provider and amqp-ping consumer
lavinmq becomes a provider of a user-facing amqp interface: a consumer
that requires a message queue is given its OWN broker — a scoped vhost
and user on a lavinmq provider — not an account on the mesh's own
control-plane broker (ADR 0048). Vhost-per-login is the isolation model,
the exact analog of postgres's database-per-login: the provider names a
vhost after the consumer's login and a user with full rights on that
vhost and none elsewhere, so a login is a broker the consumer alone can
reach.

The provider drives lavinmq through its HTTP management API (client.ts,
the module's one impure seam), with a run-once bootstrap that computes
the RabbitMQ-compatible password hash lavinmq's config wants from the
plain admin secret the mesh mints — the value no ${secret:...}
placeholder can produce and the reason the bootstrap exists (ADR 0052).
serves.amqp carries the port so consumers reference ${bound:amqp:port}.

amqp-ping is a demo consumer: it contributes nothing (the vhost is the
login), reads its grant from an env-file the mesh fills, and uses
${bound:amqp:as} for BOTH its username and its vhost — the db-name
lesson applied to AMQP. It speaks AMQP 0-9-1 over a raw socket with no
npm dependency (the way redis speaks RESP) and round-trips one message.
It carries a slug so its identity fits the 20-char backend bound
(ADR 0049).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 15:50:19 +02:00
jschoubben 89254ef2c8 hello-web: add slug so its identity fits the 20-char backend bound (ADR 0049)
mesh_anchor_hello_web is 21 chars, one over the S3 access-key bound; slug 'hello' brings it to 17.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 15:16:56 +02:00
jschoubben 36640f68e6 Add route-proxy provider and hello-web consumer modules
route-proxy is the shipping form of the reference reverse proxy (novox/hq
ADR 0007, 08-connectivity section 3): it provides route, is given every
consumer as the file at receives.route, and forwards by the Host header. It
ships the Go proxy from mesh-control/examples/route-proxy via a multi-stage
Dockerfile; no broker, own-secret or provisioner, since it only reads the file
the mesh writes.

ACME_DIRECTORY defaults to Let's Encrypt staging and is overridable per node to
production, so there is no hardcoded production default -- resolving novox/hq
04-ISSUES/004. hello-web is a minimal consumer that requires route and
contributes name+port, to exercise the grant.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 15:07:16 +02:00
jschoubben d301e05794 fix: postgres consumers connect to the db the mesh named (the login), not a hardcoded app name
The postgres provisioner creates each consumer a database named after the login the mesh
minted (mesh_<node>_<module>), per ADR 0048 -- 'a same-named database under exactly that
login'. But baserow, letta, umami, gitea, keycloak and nextcloud each hardcoded their app db
name (DATABASE_NAME=baserow, /letta, /umami, NAME=gitea, /keycloak, POSTGRES_DB=nextcloud),
so the service connected to a database that does not exist ('database letta does not exist').
Each now uses ${bound:postgres-database:as} as the db name -- the login, which is also the db
name -- matching the working meshboard pattern and the provisioner's actual behaviour.

Found by the two-node DB-consumer lab install (mesh-lab assigned-two-node-db); baserow is
proven connecting and running there. The S3 bucket name is the same class of assumption and is
a separate follow-up (s3 identity has its own length bound).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 13:47:16 +02:00
jschoubben 1f0c3ac69e fix: db/cache/store providers must serve their port; drop baserow's empty redis contribution
The breadth install (catalogue-broad) surfaced two latent resolution bugs that
single-module compile checks never caught (compile != resolve):

1. postgres/redis/mongodb/minio served no "port", yet seven consumers
   (baserow, letta, invoicing, gitea, umami, keycloak, nextcloud) reference
   ${bound:<provision>:port}. A binding auto-carries at/from/as; the port is the
   provider's half of the answer and must be declared in `serves` (manifest.go:
   "Serves is what a consumer needs to know ... a port, a path, a realm"). Added
   port to each: postgres 5432, redis 6379, mongodb 27017, minio 9000. Without it
   no DB consumer could resolve, let alone deploy.

2. baserow declared an empty contribution `contributes: {"redis-cache": {}}`.
   redis's serves spec for the cache is empty (a cache takes no per-consumer
   payload), so a consumer only `requires` it; an empty contribution is refused.
   Dropped it (requires/binds/secrets unchanged).

Found by mesh-lab catalogue-broad; the same fixes are mirrored in that bed.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 02:38:22 +02:00
jschoubben db8b60f4e4 Convert confluence and jira into nox catalog modules
Port the HAL confluence and jira integrations to the tools-only,
outbound-only external-SaaS pattern proven by the merged gitlab module:
runtime-only containers (container-runtime capability), own-secret token +
broker, a settings-managed config.json for public config, and a
per-module runtime image.

Each module carries its own Atlassian API client and tools (ADR 0039),
translated from HAL's @hal/sdk zod-schema/MCP-content shape into
mesh-sdk's input/run-returns-data shape. Public config (ATLASSIAN_URL,
ATLASSIAN_EMAIL) lives in config.json; the API token is the one
own-secret. Clients are built lazily and never throw at registration, so
each runtime serves its full tool surface with no credentials (the
Servarr lesson) — confluence serves 3 tools, jira serves 8.

jira's periodic ticket-poller (update-tickets.service/.timer) is NOT
ported: the mesh has no scheduled-task primitive yet (a pending
decision). Only jira's tools are ported; a top-of-file note records the
deferral.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 01:49:22 +02:00
jschoubben 7eb156a82a Convert gitlab into a tools-only mesh catalog module
Port the HAL gitlab module (whose tools lived in @hal/sdk) into a
self-contained mesh-catalog module modelled on cloudflare-dns: the GitLab
API client and all its tools live in the module (ADR 0039), served through
mesh-sdk's registerModuleTools harness.

Tools-only, outbound-only external-SaaS shape: a runtime-only container on
network:host, no service, no listener, no provisioner. The token is an
own-secret; GITLAB_URL is a public setting in a merge:json config file.

The client is built lazily and never throws at registration, so the runtime
comes up and serves all 23 tools even with no valid token (the Servarr
lesson) — it only fails when a tool is actually invoked unconfigured.

Ported 23 tools: projects (list, get), merge requests (list, get, create,
approve, add note), pipelines (list, get, retry, cancel, list jobs, job log),
and project + group CI/CD variables (list, get, create, update, delete each).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 01:33:43 +02:00
jschoubben 8b82e47889 invoicing: point app/api at the real registry host
The conversion left the image host as a placeholder (registry-api.example).
Point both containers at the actual registry the built app/api images live in
(novox/invoicing-{app,api}, tags 51-55 + latest present). Tracks :latest to
match the app's continuous-deploy model; pin by digest later if reproducibility
of a specific build is wanted.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 01:14:08 +02:00