Commit Graph
173 Commits
Author SHA1 Message Date
jschoubben 071afc1e2c Merge pull request 'Modules compile in the toolchain and ship on the runtime' (#22) from feat/build-and-run-are-two-images into main 2026-09-14 02:02:42 +02:00
jschoubben cf116a4932 Modules compile in the toolchain and ship on the runtime 2026-09-14 01:57:06 +02:00
jschoubben 5ea88b149b The three modules name their base rather than pinning a copy of it
Each named a digest produced inside a lab that no longer exists, so none of them
could be built anywhere else. They say which module they stand on now, and there
is deliberately no default — a build nobody told stops at the declaration rather
than at a reference that resolves to nothing.
2026-09-13 23:53:22 +02:00
jschoubben 720e3706b2 Merge pull request 'The catalogue holds the module graph, and the modules that make it build themselves' (#21) from feat/the-catalogue-module into main 2026-09-13 11:17:10 +02:00
jschoubben 729582cc55 Remove what was only there to move a commit
A line in postgres's recipe and a one-line file in amqp-ping, both added to
move a commit and watch the mesh notice. The proofs worked; neither was meant
to stay. MESH_MODULE is set from the sealed credential at run time anyway, so
baking it in was dead weight as well as noise.
2026-09-13 11:15:33 +02:00
jschoubben 21ad008879 Stale means the artifact moved, not the commit
A comment changed in a build recipe is a new commit and a byte-identical image.
Comparing commits called every module standing on it stale, so the mesh would
have rebuilt itself entirely to arrive back exactly where it started — and
listed each dependent once per commit that had produced the same image.
2026-09-13 02:48:21 +02:00
jschoubben 87243bc524 Every module stands on a base the mesh built
The base was a digest typed in by hand, for an image nothing in the mesh could
produce — so the graph held edges pointing at it with no version on the far end,
and the one change that reaches every module at once could never be noticed.
It is a module now, and these edges resolve.
2026-09-13 02:45:37 +02:00
jschoubben 1d0d9a3894 Name the module in its own runtime, and move the artifact with it 2026-09-13 01:58:48 +02:00
jschoubben 358c7d5a2d Move postgres's commit, to watch the mesh roll it out by itself 2026-09-13 01:57:00 +02:00
jschoubben d5300120d6 Move amqp-ping's commit, to see what the catalogue announces 2026-09-13 01:49:48 +02:00
jschoubben c80d9d0f7c The graph holds what a module declares, not only what it was built on
Build edges are discovered by building; requires and provides are stated by the
module about itself. Both belong in the graph and answer different questions —
and "what provides postgres-database" needed a sweep over every manifest, which
only something holding all of them can do.
2026-09-13 01:44:56 +02:00
jschoubben 2d2ca80e3b Index the edge after the column that carries it exists
On a store with the old shape the table is not re-created, so an index declared
beside it is built on a column the migration has not added yet.
2026-09-13 01:41:23 +02:00
jschoubben f15c814145 A build edge names an artifact, because that is what the builder can see
The catalogue expected each edge to name a module and a commit. The builder
sends a pinned image reference — it cannot know which module produced it, that
is a fact about the graph. So every build that had been built on top of anything
was rejected, and only the modules built against nothing ever registered.

Versions now record what they published, and an edge resolves through that. An
edge to an artifact no module here produced is kept: it resolves by itself when
that module is registered, which is the ordinary case while a mesh fills in.
2026-09-13 01:37:51 +02:00
jschoubben 9387f8b040 A module that listens is served, not run
`run` imports an entrypoint without binding a broker — it exists for a step that
works offline and exits. Both the catalogue and amqp-ping subscribe on import,
so both died on the first on() with no broker bound.
2026-09-13 01:30:39 +02:00
jschoubben c774d5dbe0 Install the postgres client the way the runtime base can
The published base is debian; apk is not there and the build said so.
2026-09-13 01:28:15 +02:00
jschoubben 6ebf51d312 postgres's runtime carries the client it provisions through
Its provisioner runs DDL by shelling out to psql, which the runtime base has no
reason to hold. Every create failed with ENOENT and retried for ever.
2026-09-13 01:26:23 +02:00
jschoubben 594295f009 The catalogue restarts when its database credentials change
Without it the container keeps whatever the env file said when it was created.
Nothing reports that: it runs, and it is wrong.
2026-09-13 01:22:38 +02:00
jschoubben f7d57e9556 The catalogue brings its own postgres driver
The runtime base carries what every module needs, and a database driver is not
that. Installed into an empty directory because the module's package.json also
names the sdk, which lives in the base rather than on a registry.
2026-09-13 01:16:23 +02:00
jschoubben e742b6a569 The catalogue builds its own runtime, and postgres serves its tools
Both modules keep their tools in an entrypoint of their own, so an image that
named only the consumer would serve none of them.
2026-09-13 01:13:43 +02:00
jschoubben d6c9c8d666 postgres builds its own runtime, like any other module
Its provisioner container named an image nobody could produce — a zero digest
placeholder. It names an artifact instead, and the module says how to build it,
so the mesh can make the database provider the catalogue needs.
2026-09-13 01:10:45 +02:00
jschoubben 2a6fed6f4f The builder is told the mesh's name for the machine it runs on 2026-09-13 01:07:37 +02:00
jschoubben 81a80c675c The builder declares what it announces
Its account is scoped from what it emits and consumes, and it declared neither —
which is why asking for a generic module account produced one that authenticated
and could do nothing, with the refusal surfacing a layer away as a permissions
error against a queue.

Declaring the announcement is not documentation here. It is what the permission
is derived from.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-13 00:57:41 +02:00
jschoubben 39631d6f87 amqp-ping says what it is made of, and can be built from its own directory
A module's runtime image was assembled by a script copying the sdk and the tool
runtime out of neighbouring checkouts, so it could only be built on a workstation
that had them. That is why no module declared what it was made of and why
forty-seven point at a placeholder.

The tool runtime becomes an image a module's runtime is built FROM, published like
any other artifact. The module then builds from its own directory and that base —
one clone, which is what the builder can actually be asked for (novox/hq ADR 0069).
The dependency stops being a property of somebody's machine and becomes a build
edge, pinned to a digest the mesh's registry assigned.

The compiler is invoked by its real path rather than through node_modules/.bin:
those are symlinks to a launcher that requires its library relatively, and
resolving them while building the base leaves a launcher pointing at nothing.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-13 00:52:57 +02:00
jschoubben a73cb8a2f8 The catalogue, as a module that owns the module graph
It links module-versions to each other and knows nothing about nodes; which
machine runs what stays the control plane's (novox/hq ADR 0070, 0072). Keeping
them apart is what lets the control plane carry on composing declarations while
this is down.

The builder announces what it built, this places it in the graph and announces
what that means, and the control plane hooks the meaning rather than the build
output. A rebuild producing the commit already current is registered and is not
an upgrade — announcing it would ripple outward forever through modules that did
not change.

Ordering is not computed. Modules stale and waiting on nothing that is itself
stale are announced as buildable; the rest stay stale and appear once whatever
they were waiting for is registered, so a chain and a diamond need no special
handling and nothing holds a plan.

Four tools over the graph: what this mesh holds, one module in full, what a
change to a module reaches, and what must be rebuilt and why. The edges are
derived from builds rather than declared, so they cannot drift from what the code
actually uses.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 23:16:05 +02:00
jschoubben a4d10341e7 Declare what hello-web is made of, to prove a build from a path
A proof branch, not for main: the route-forwarding bed reads this module's literal
image and would break until the module is built.

The modelling is right regardless — the image is upstream, so the mesh should
mirror it once into its own registry and pin what that registry assigned, rather
than every machine fetching a reference somebody else can move.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:57:18 +02:00
jschoubben 5f76e34994 Add the builder as a module, so the mesh can be given one
The code that turns a repository into artifacts already existed and is deliberately
something a machine runs as an ordinary module rather than something the control
plane does. There was no module for it, so it could never be placed and never ran —
which is why nothing in the catalogue could be produced.

It asks for a container runtime because it builds images, requires the artifact
store because it publishes into it, and claims one per machine. Its credential
folder is mounted rather than the file, since a file mount keeps pointing at the
old contents after the mesh writes new ones.

The store is reached on the machine's own loopback: the address in a binding is
where a machine sits on the private network, and the store has to be co-located
anyway because runtimes refuse a plain-HTTP registry anywhere else.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:45:59 +02:00
jschoubben 5aacd2538b mesh-control: the enrolment endpoint is the substrate's, not the overlay's
The module said MESH_BROKER_ADDRESS=${machine:at}:5671, and that cannot work in either
direction.

At genesis it does not resolve at all. `at` is a machine's name on the PRIVATE network,
and the mesh only holds one for a node that has an overlay placement and resolves the
networking module — neither of which exists when the control plane is installed, which
is step 9 of ten, long before anything has been placed anywhere. mesh-control refuses a
${machine:} key it does not hold rather than writing the literal through, so the push
would have stopped with "this machine says name".

And afterwards it would be the wrong address anyway. This value is what every enrolment
token tells a joining node to dial. A machine that has not enrolled is not on the
overlay, so an overlay name is precisely the one thing it cannot reach.

It is the substrate's own fact — the address the broker advertises, decided by whoever
wrote the bundle, which the mesh did not make and cannot invent. So it arrives the way
the store connections beside it arrive: an own-secret the installer delivers with
`secret accept`, read out of the bundle it produced. mesh-bootstrap already does this for
every variable the module fills from a secret; this one simply joins them.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 11:48:51 +02:00
jschoubben 2c322cb2fb mesh-control: the control plane could not read its own connections
Its image is FROM scratch and runs as 65534. The host writes a sealed own-secret 0600,
owned by root, which is right — but the module then bind-mounted those three files into
the container and told the process to open them. It cannot:

  $ docker run --rm -v <0600 root file>:/run/secrets/inventory:ro \
      -e MESH_STORE_INVENTORY_FILE=/run/secrets/inventory mesh-control:development status
  MESH_STORE_INVENTORY_FILE names /run/secrets/inventory ... and it cannot be read:
  open /run/secrets/inventory: permission denied

Measured on a workstation, not reasoned about. Every other module in this catalogue gets
away with the same mount because its runtime container runs as root; this one does not,
and genesis (novox/hq ADR 0067) would have stopped at step 9 with a control-plane module
that starts and cannot open a context.

The connections go through the env file this module already has instead. That file is
mode 0600 and is read by the container runtime's client, which is root — the same reason
the broker's URL has always reached the process this way. It also sidesteps the inode
that a file bind mount pins (290be37, step-ca): --env-file is read afresh at create, and
restart-on names it.

The own-secrets stay exactly as they were, because the installer delivers the substrate's
real connection strings into them with `secret accept` before the first push — the mesh
did not make those credentials and cannot invent them.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 11:38:28 +02:00
jschoubben 290be37a93 step-ca: mount the directory, not each secret file
A file bind mount tracks the inode. The host writes atomically — new file, rename
over — so the container keeps reading the file that was there when it started,
and a rotated or newly-delivered secret never reaches it. Mounting the parent
directory resolves the path on each open instead.

This is a known shape (hal KB troubleshooting/docker-bind-mounts), and it cost an
hour here before it was looked up: the CA crash-looped on a root key it had
already been given, because the container still held the inode from before the
key arrived.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 08:49:55 +02:00
jschoubben 0d2c2dd989 minio: pin the manifest these machines can actually run
The pinned digest named the arm64 manifest, so an x86 machine pulled it and the
container died with 'exec format error' on every restart. It never showed while
the lab ran its own registry: the harness pushed the WORKSTATION's copy, which
is amd64, and every machine then pulled that under a digest the registry had
just assigned. The registry was quietly correcting the architecture too.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 02:51:34 +02:00
jschoubben 9ff320ecca Pin the operator's own images by digest, as every other image already is
Nine references across seven modules named ':latest'. ADR 0006 forbids it and
the host refuses it by name — and the refusal had never fired, because the lab
pushed every image into its own registry and rewrote each reference to the
digest it had just assigned. Deleting that registry made these the only
manifests the host would now reject (novox/hq 04-ISSUES/039).

The digests are what each tag resolves to today, read from the registry that
serves them.

This is a stopgap and should be said as one: a digest written into a repository
is wrong the moment anybody rebuilds, which is precisely why the design has the
repository name artifacts and the mesh hold digests. Until something builds and
publishes, a digest that is stale is still better than a tag that silently moves.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 01:06:14 +02:00
jschoubben 4a1c5ac817 The artifact store requires nothing, because at genesis there is nothing
A route-label migration gave the registry a public name, and with it a
requirement. But the registry is the first module a new mesh installs: at that
moment nothing provides a route, so the push is refused and a mesh cannot get
its own image store — the cycle 04-ISSUES/029 closed, re-entered through a
different door.

The same rule that record states applies: a module providing the artifact store
cannot depend on what the store is needed to deliver. A public name for it is an
ordinary want and belongs to a module beside it, installed once there is a mesh
to install things.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:06:23 +02:00
jschoubben 71f8012c5b The control plane as an ordinary module
novox/hq ADR 0067 pivots genesis through a temporary control plane and then
reinstalls the control plane as an ordinary module pinned to a digest the mesh's
own registry assigned. That record notes the one thing missing: a control-plane
module manifest, which did not exist.

It could not be written honestly before now. The control plane read its store
connection from MESH_STORE_<CONTEXT>, that connection string carries a password,
and a manifest can put a sealed value into a file's `content` but has nothing
that substitutes into a container's `env`. So the manifest could carry the
password in the clear, or omit the setting. mesh-control now also accepts
MESH_STORE_<CONTEXT>_FILE, which is how every other module here is given secret
material, and the manifest follows.

What the substrate bundle gives the control-plane container today, and where
each part has gone:

  MESH_STORE_INVENTORY   own-secret `inventory`, mounted, named by _FILE
  MESH_STORE_IDENTITY    own-secret `identity`,  mounted, named by _FILE
  MESH_STORE_LICENCES    own-secret `licences`,  mounted, named by _FILE
  MESH_BROKER_AMQP       own-secret `broker`, through an env-file hole
  MESH_BROKER_MANAGEMENT own-secret `broker-management`, likewise
  MESH_BROKER_ADDRESS    ${machine:at}:5671 in that same env-file
  MESH_BROKER_CERTIFICATE  plain env; the path is not a secret
  network host, args ["serve"], the broker's TLS volume  unchanged

The two broker URLs go through an env-file rather than a file of their own
because mesh-control has no MESH_BROKER_AMQP_FILE. That is the same fault one
layer over, and the same remedy would fix it; it is out of this change's scope
and is written down rather than papered over.

None of these values is in the manifest. Each is an own-secret the operator
supplies with `secret accept` — the mesh cannot invent a connection string — and
the container restarts when any of them changes.

The module claims `the-control-plane` at mesh scope, which the bundle has no way
to say: two control planes writing one inventory is a fault worth refusing at
assignment. It carries no `listens`, because `serve` dials the broker and binds
nothing. The image is the catalogue's placeholder digest for a mesh-built image,
which the installer replaces with what the registry assigned.

Checked with the real parser: all 67 manifests through catalogue.ParseManifest
and every module's CheckIdentity against all four node names — 0 problems — and
this manifest rendered through Resolution.Declaration, so the ${secret:…} names,
${machine:at}, the restart-on ids and the image pin are exercised rather than
merely parsed. Slug `control`: mesh_shanks_control is 19 of the 20 an S3 access
key keeps.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:43:07 +02:00
jschoubben 16b4dc9ab9 step-ca: certify the machine the mesh reaches it at
Its API certificate carried localhost only, so a proxy dialling the address the
mesh handed over refused it on hostname verification. The names now compose from
the machine the module was assigned to, which a manifest could not know and now
does not have to (mesh-control ${machine:...}).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:12:48 +02:00
jschoubben d9fbc72565 step-ca: give the CA's own secret material to the uid the CA runs as
The upstream `smallstep/step-ca` image runs as uid 1000. An own-secret lands as
a file the host writes root:root 0600 — the module says only a name and a path,
so there is nowhere to say who must be able to read it — and the container
crash-looped on `permission denied` reading its own root key. The lab got past
it first with `chmod 0644`, which hands the root key to every local user, and
then by running the CA as root, which is worse.

Neither is needed. A module composes its own files, and a `file` resource takes
both a `mode` and an `owner`, with `${secret:name}` reaching the module's own
secrets — the mechanism redis already uses to hand its password to a server
running as 999. So the three pieces of init material are declared as owned
files: still 0600, owned by 1000:1000, and those are what the container mounts.

The raw own-secret files stay where they were. Nothing mounts them now; they are
how the secret comes to exist on the machine, and the host is the only thing
that reads them. Same shape as redis's `default.secret`.

Verified against the real code rather than by inspection: mesh-control attaches
the sealed value to each of the three resources and keeps `owner` and `mode`
(`sealedFor` + `intoFile` over this manifest), and mesh-host parses `owner` on a
sealed-substituted file and lands it 0600 owned by 1000:1000 (`applyFile`).
2026-09-10 20:57:58 +02:00
jschoubben 43ca9c9c73 Give every module that would overflow its login a short slug
An identity is `mesh_<node>_<slug-or-name>` and a backend keeps 20 characters
(an S3 access key). Overflow makes a module unresolvable, and this catalogue
was finding it one module at a time, on a raise: route-proxy on novox is 22,
home-assistant on ace is 23. Two found by hand where a sweep would have found
eighteen.

So the whole catalogue was swept instead, against the longest node name the
mesh actually has (`shanks`, six characters) rather than against the node each
module happens to sit on today — a module is assigned somewhere, and where is
not a property of the manifest. That leaves eight characters for the identity
source, and eighteen modules were over it.

Slugs added, chosen to stay greppable in a provider's user list:

  anthropic-consumer  claude     openai-consumer     openai
  anthropic-manager   anthmgr    portainer           portain
  audit-logger        audit      public-acme         pubacme
  bookshelf           books      qbittorrent         qbt
  cloudflare-dns      cfdns      resolv-conf         resolv
  confluence          confl      resolved-split-dns  splitdns
  home-assistant      hass       route-proxy         rproxy
  invoicing           invoice    verdaccio           verdacc
  mosquitto           mosq
  nextcloud           ncloud

A slug changes the login the mesh mints, so a module already provisioned under
its full name is re-minted under the slug and its old login withdrawn — which
is the provisioner's ordinary business, but it is a change, not a no-op.

Checked with the real parser: every one of the 66 manifests through
`catalogue.ParseManifest`, and every module's `CheckIdentity` against all four
node names. 0 problems, where the same check over the parent commit reports 44.
2026-09-10 20:57:45 +02:00
jschoubben 08be6683b6 route labels: nextcloud is 'drive' (real hostname), novox.be is the apex '@'
nextcloud's label was migrated from a wrong module-name default; its real
production hostname is drive.novox.be. novox.be is the bare-domain apex, now the
'@' label (composeName gained apex support).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 00:27:25 +02:00
jschoubben f0665ba956 ADR 0056: selectable ACME issuer, step-ca init from operator root, route labels
Piece A + B of ADR 0056, completing the internal-CA work in ff01ada.

Selectable issuer. acme-ca is now a role two providers can satisfy: step-ca
(internal CA) or the new public-acme (a fact-only module, no container/listen)
that serves Let's Encrypt production. A mesh assigns one or the other to satisfy
route-proxy's `requires: acme-ca`.

One directory shape for both. A provider serves the ACME directory's parts the
way the mesh already models any reachable service -- an address (`at`), a `port`
and a `path` -- and route-proxy composes `https://<at>:<port><path>`. step-ca
lets the mesh fill `at` (its node) and `port` (its single listen) and serves only
`path`; public-acme, not being a mesh service, serves all three (overriding `at`
with the public host). Same composition either way.

Empty root means the system trust store. Both providers serve `root`: step-ca
the operator root PEM (settled per mesh), public-acme an empty string. route-proxy
writes it to the CA bundle file unconditionally; the binary now reads an empty
bundle as "the root is already trusted by the OS" and falls back to system roots
(examples/route-proxy/main.go, committed on the mesh-control ADR-0056 branch).

step-ca inits from the operator's root. The operator's root cert, root key and
root-key password are mounted at the smallstep entrypoint's default init paths
(/run/secrets/root_ca.crt, root_ca_key, root_ca_key_password) with the matching
DOCKER_STEPCA_INIT_*_FILE vars, so `step ca init` adopts the operator's root
instead of self-generating one -- the CA that signs is the CA route-proxy trusts.

Route names are labels, not FQDNs. Every routed module now contributes a `label`
(the leftmost subdomain) instead of a full public hostname; the node's public
domain composes the name. Apex (novox.be) is left as a full name -- composeName
has no empty-label/apex convention yet (mesh-control follow-up).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 00:21:42 +02:00
jschoubben ff01adaa45 Add internal ACME CA (step-ca) as a generic acme-ca provision; wire route-proxy to consume it
Piece C of ADR 0056: an internal authority certifies routed names by the same
path a public one would, with the proxy pointed at whichever issuer the mesh
names and trusting that issuer's root.

step-ca (new): provides acme-ca (mesh scope), listens 9000 from mesh, persists
its CA home under uid 1000. The root CA cert reaches consumers as a served
value settled from a per-mesh operator setting (no baked root); the init
password and root key are sealed secrets, not literals. No route-name -> IP
hosts map (that is Piece B). Upstream smallstep/step-ca pinned by Docker Hub
digest.

route-proxy: requires + binds acme-ca. ACME_DIRECTORY is composed on the
consumer side from ${bound:acme-ca:at}:${bound:acme-ca:port}, and ACME_CA_BUNDLE
is a file whose content is ${bound:acme-ca:root} -- both from the binding,
replacing the hardcoded directory and the baked root PEM.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 23:07:54 +02:00
jschoubben d018f1e16e Route names mirror real production hostnames from live Traefik
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 14:04:17 +02:00
jschoubben e26ca38eaa novox conversions: give the web apps public names through route-proxy
Add a `route` contribution (requires/contributes/binds) to every web app so
each gets a Host-routed public name via route-proxy, mirroring the
de-spiegel/only-office pattern:

- novox: gitea, keycloak, nextcloud, umami, invoicing, verdaccio, registry,
  and mailu (single mail.novox.be -> 7080; admin/webmail/api ride that port).
- ace: grafana, sonarr, radarr, lidarr, bazarr, ombi, tautulli, jackett,
  nodered, searxng, home-assistant, bookshelf, baserow.

Split photos so its three sites each get a name: photos keeps server +
admin-client (photos.novox.be), and new photos-eef (eef.novox.be) and
photos-filip (filip.novox.be) modules carry the client sites.

The production FQDN stays the literal default; a per-node .incus name is a
settings override applied where the mesh runs, not a manifest hardcoding.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 12:51:15 +02:00
jschoubben 431310fb03 novox conversions: short slugs for the login-length cap + real roundcube digest
The whole-mesh-novox lab validation found three modules unresolvable: a
provision-consuming module's minted login is mesh_<node>_<module>, and on novox
`mesh_novox_only_office`(22)/`_de_spiegel`(21)/`_amqp_email_forwarder`(31)
overflow the 20-char S3 access-key cap (ADR 0049). Added a short `slug` each
(office/spiegel/emailfwd → 17/18/19 chars). Also fixed mailu's roundcube image:
the pinned digest returns "manifest unknown" from ghcr; corrected to the real
:1.9 digest.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 02:35:47 +02:00
jschoubben b3309a0ce9 mailu: full multi-container stack from the live novox deployment
Rebuild the mailu manifest from the running production deployment as the
source of truth, and finish wiring its tool runtime.

- Full 11-container topology: front, smtp, imap, admin, antispam, antivirus,
  webmail, webdav, fetchmail, resolver, redis — real ghcr.io/mailu images
  pinned by digest at tag 1.9 (clamav/radicale/fetchmail digests newly fetched).
- Admin DB now consumes the mesh postgres-database provider (requires +
  contributes + binds + secrets, DB_* templated from ${bound}/${secret}),
  replacing the bundled postgres:13 admindb the live stack still runs.
- Full config env from the live containers as a plain env-file; SECRET_KEY,
  the admin API token and the initial-admin password become own-secrets;
  DB password comes from the provider secret. No secret values hardcoded.
- Enable the Mailu admin REST API (API=true, WEB_API, API_TOKEN own-secret) so
  the ported tools can reach it — the live deployment runs this API OFF.
- mesh-mailu tool runtime on the mailu network: admin API over the module
  network, token from the mounted own-secret, docker.sock for the doveadm mail
  reads, broker + mergeable config with restart-on.
- client.ts fromEnv reads the API token from its mounted own-secret file
  (MESH_MAILU_API_KEY_FILE), matching the cloudflare-dns/umami pattern.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 01:04:55 +02:00
jschoubben 8f99b5bd42 Add amqp-email-forwarder daemon module
Owned daemon that consumes from the mesh amqp provision (lavinmq) and
forwards to SMTP. No ports, no route. AMQP host/port/user/password
templated from the amqp grant (analog: amqp-ping); vhost/exchange/queue
are the app's own config; SMTP user+password are given own-secrets
(secret accept). See report for the EMAILDELIVERY_T vhost integration flag.
2026-09-09 00:41:06 +02:00
jschoubben c423e21a8e Add novox.be; replace photos immich stub with the real app
- novox.be: owned www:latest container, public route novox.be
  (host 4000 -> app 8080). No DB, no secrets. Analog: de-spiegel/hello-web.
- photos: the existing module.json was a wrong immich stub (alpine image,
  port 2283). Replace it with the user's real photo app: photos-server
  backend consuming the mesh s3-bucket (bucket photos) + mongodb-database
  (db photos) providers, env templated from the provisions (analog:
  invoicing), plus the three static client containers (admin + two family
  sites). Only photos.novox.be is routed: contributes.route is single-valued
  across the catalog, so eef/filip need their own modules (flagged).
2026-09-09 00:36:48 +02:00
jschoubben 1fc10be11b Convert only-office, n8n, de-spiegel to nox modules
Mirror the proven converted modules: pure manifests, no runtime.

- only-office: owned onlyoffice/documentserver container, generated JWT
  own-secret, public route office.novox.be. Analog: gitea/hello-web.
- n8n: consumes mesh postgres-database + redis-cache providers instead of
  bundling them, owned container, public route n8n.novox.be. Analog:
  baserow (postgres+redis) + hello-web (route).
- de-spiegel: owned container built from its own repo image, public route
  de-spiegel.novox.be, given SMTP own-secrets. Analog: hello-web/invoicing.
2026-09-09 00:30:06 +02:00
jschoubben 6a84e97e53 Merge pull request 'module fixes from the whole-mesh dry-run: fail2ban capability + tool-runtime credential wiring' (#20) from feat/module-cred-fixes into main 2026-09-08 18:45:24 +02:00
jschoubben 5973d41966 umami: read the admin password from its mounted secret file
The whole-mesh dry-run found umami's runtime crash-looping "admin password is
not set": its `admin` own-secret is mounted at /run/secrets/admin, but the
client read the bare env UMAMI_ADMIN_PASSWORD, which nothing sets. Same shape as
the six tool-runtime credential fixes — read the mounted file first
(MESH_UMAMI_ADMIN_PASSWORD_FILE), falling back to the env. (photos and mailu
remain deeper conversion jobs — a stub app image and a full Mailu config env —
not credential-wiring, tracked separately.)

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-08 18:43:48 +02:00
jschoubben d289e5a928 modules: wire tool-runtime app credentials as own-secrets
The six modules that run a mesh-<mod> tool-runtime sidecar read an app
credential from an env var the manifest never provided, so the sidecar
crash-looped in the whole-mesh dry-run (e.g. "no Plex token — set
MESH_PLEX_TOKEN"). These are operator-set app secrets, so deliver them the
same way cloudflare-dns delivers its API token: an own-secret file mounted
read-only, with a MESH_<APP>_*_FILE env pointing at the mount, and the
runtime code preferring that file (falling back to the existing env so
nothing regresses).

- plex: own-secret token -> /run/secrets/token, MESH_PLEX_TOKEN_FILE
- bazarr: own-secret api-key -> /run/secrets/api-key, MESH_BAZARR_API_KEY_FILE
- ombi: own-secret api-key -> /run/secrets/api-key, MESH_OMBI_API_KEY_FILE
- home-assistant: own-secret token -> /run/secrets/token, MESH_HOMEASSISTANT_TOKEN_FILE
- nzbget: own-secret password -> /run/secrets/password, MESH_NZBGET_PASSWORD_FILE (URL stays plain env)
- qbittorrent: own-secret password -> /run/secrets/password, MESH_QBITTORRENT_PASSWORD_FILE (URL stays plain env)

The operator now completes each with `secret accept <node> <module> <name> --from <file>`.
tsc passes for all six.
2026-09-08 18:40:23 +02:00
jschoubben 75fb16bbfb fail2ban: require the firewall capability, not the non-existent intrusion-prevention
The whole-mesh dry-run found fail2ban unassignable on every node: it declared
`capabilities: ["intrusion-prevention"]`, which mesh-host has no detector for
(its detectors are container-runtime, package-manager, service-manager,
firewall, overlay, graphical-session, seat, privileged). intrusion-prevention
is what fail2ban PROVIDES, not a host capability it needs. It bans via
iptables/ufw, so it needs `firewall` — the same capability the firewall module
declares. The `the-intrusion-prevention` claim (node-exclusive) is unchanged.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-08 18:27:34 +02:00