OBJECTSTORE_S3_BUCKET=nextcloud was a leftover from before the module
existed — that bucket was created by hand during tonight's earlier HAL
credential stopgap. The mesh's own minio provisioner derives its own
bucket name from the consumer's access-key identity (bucketFor(as) in
minio/client.ts) rather than honouring contributes.s3-bucket.bucket — by
design, so teardown can recompute the name with nothing persisted — and
minted mesh-novox-ncloud, a different bucket. mesh_novox_ncloud's scoped
policy only covers that bucket, so every S3 write 403'd with AccessDenied
trying to touch the old one. Pointed both the request hint and the real
env var at the bucket that's actually there.
the s3-bucket binding already carries serves.region (same mechanism as
at/port); ${bound:s3-bucket:region} tracks whatever minio is actually
configured with instead of a copy that can drift
minio runs with MINIO_REGION=eu-west; nextcloud's S3 config never set a
region, so every object write (avatars, file writes) failed signature
validation with AuthorizationHeaderMalformed, surfacing as Internal Server
Error on real page loads
the migrated data has no literal 'admin' account — HAL's real admin login
is a personal account (jochens), not a generic one. Resetting that would
touch a real user's own credential, so the module gets its own dedicated
admin-group service account instead, same pattern as the minio per-module
service accounts. mesh-admin was created once by hand on novox to match
this manifest for the already-migrated data; a genuinely fresh install
seeds it automatically via NEXTCLOUD_ADMIN_USER/NEXTCLOUD_ADMIN_PASSWORD.
build refused to reach docker:cli implicitly (novox/hq ADR 0097); pin it by
digest and thread it through as DOCKER_CLI, redeclared in the final stage
since args declared before the first FROM don't carry past it
- MESH_NEXTCLOUD_URL hardcoded :80 instead of the mesh-assigned ${port:80}
- sidecar's occ() shells to docker exec but the docker CLI binary was never
present in the runtime image, only the mounted socket
The mesh's own check caught it: an env-file-loaded secret still reaches
the process environment, readable via docker inspect and /proc (hq
04-ISSUES/041) -- the same class of exposure the file-based delivery
exists to avoid. Added MESH_NEXTCLOUD_ADMIN_PASSWORD_FILE support to the
client, matching the pattern the minio client already uses, and mounted
the sealed admin secret directly rather than writing it into an env-file.
The sidecar's own client needs MESH_NEXTCLOUD_ADMIN_PASSWORD to list
shares over the OCS API, but the runtime container's env/volumes never
carried it -- only the server container did. Delivered the same way every
other sealed value in this manifest already is: a generated env-file with
the ${secret:admin} substitution, not a raw value in the container's env.
The pinned digest resolved to a PHP 8.5.10 image; Nextcloud 30 refuses to
run above PHP 8.4. Repinned to the current digest for the nextcloud:30 tag
(matches HAL's own NEXTCLOUD_VERSION), which carries PHP 8.3.28 -- the
same version the data being migrated was actually running under.
From the survey of every env-file secret (ADR 0086, issue 041): amqp-ping,
minio, mongodb and grafana use the _FILE twin their software honours;
mesh-catalog and model-usage read DATABASE_URL_FILE (a file the mesh
templates, mounted where only the runtime reads it); grafana's secret files
belong to its own account. Two dead deliveries removed: a line nothing read
in amqp-email-forwarder, and mailu's secret.env on four containers that
never read it. The 25 exceptions that remain carry the surveyed reason —
convertible and awaiting a bed, convertible through a generated config file,
the application's own code, or not convertible.
ADR 0086. mesh-controller mounts its six own secrets and names them with
_FILE twins, so no credential of its own reaches its environment. The 35
containers that still read a secret through an env-file carry
secrets-in-environment with the reason; converting each where its software
accepts a path is the per-module work of issue 041.
keycloak, mailu, minio, mongodb, mssql, nextcloud, portainer, redis,
umami and verdaccio get the Dockerfile + build section the eight
buildable modules already had; their runtime containers name the
artifact instead of a placeholder digest.
One convention, settled (060's open question, informed by 061): the
runtime container runs serve mode with every serve-time entrypoint in
MESH_TOOL_MODULES — tools serve, events flow, and a provider's
provisioner reconciles in the same process with the broker connected.
postgres, gitea and lavinmq are retrofitted from args-run provisioners,
which served no tools and emitted lifecycle events nowhere.
route-proxy is deferred: its build context is the mesh-controller
repository, a cross-repo shape the build section cannot yet express.
An identity is `mesh_<node>_<slug-or-name>` and a backend keeps 20 characters
(an S3 access key). Overflow makes a module unresolvable, and this catalogue
was finding it one module at a time, on a raise: route-proxy on novox is 22,
home-assistant on ace is 23. Two found by hand where a sweep would have found
eighteen.
So the whole catalogue was swept instead, against the longest node name the
mesh actually has (`shanks`, six characters) rather than against the node each
module happens to sit on today — a module is assigned somewhere, and where is
not a property of the manifest. That leaves eight characters for the identity
source, and eighteen modules were over it.
Slugs added, chosen to stay greppable in a provider's user list:
anthropic-consumer claude openai-consumer openai
anthropic-manager anthmgr portainer portain
audit-logger audit public-acme pubacme
bookshelf books qbittorrent qbt
cloudflare-dns cfdns resolv-conf resolv
confluence confl resolved-split-dns splitdns
home-assistant hass route-proxy rproxy
invoicing invoice verdaccio verdacc
mosquitto mosq
nextcloud ncloud
A slug changes the login the mesh mints, so a module already provisioned under
its full name is re-minted under the slug and its old login withdrawn — which
is the provisioner's ordinary business, but it is a change, not a no-op.
Checked with the real parser: every one of the 66 manifests through
`catalogue.ParseManifest`, and every module's `CheckIdentity` against all four
node names. 0 problems, where the same check over the parent commit reports 44.
nextcloud's label was migrated from a wrong module-name default; its real
production hostname is drive.novox.be. novox.be is the bare-domain apex, now the
'@' label (composeName gained apex support).
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Piece A + B of ADR 0056, completing the internal-CA work in ff01ada.
Selectable issuer. acme-ca is now a role two providers can satisfy: step-ca
(internal CA) or the new public-acme (a fact-only module, no container/listen)
that serves Let's Encrypt production. A mesh assigns one or the other to satisfy
route-proxy's `requires: acme-ca`.
One directory shape for both. A provider serves the ACME directory's parts the
way the mesh already models any reachable service -- an address (`at`), a `port`
and a `path` -- and route-proxy composes `https://<at>:<port><path>`. step-ca
lets the mesh fill `at` (its node) and `port` (its single listen) and serves only
`path`; public-acme, not being a mesh service, serves all three (overriding `at`
with the public host). Same composition either way.
Empty root means the system trust store. Both providers serve `root`: step-ca
the operator root PEM (settled per mesh), public-acme an empty string. route-proxy
writes it to the CA bundle file unconditionally; the binary now reads an empty
bundle as "the root is already trusted by the OS" and falls back to system roots
(examples/route-proxy/main.go, committed on the mesh-control ADR-0056 branch).
step-ca inits from the operator's root. The operator's root cert, root key and
root-key password are mounted at the smallstep entrypoint's default init paths
(/run/secrets/root_ca.crt, root_ca_key, root_ca_key_password) with the matching
DOCKER_STEPCA_INIT_*_FILE vars, so `step ca init` adopts the operator's root
instead of self-generating one -- the CA that signs is the CA route-proxy trusts.
Route names are labels, not FQDNs. Every routed module now contributes a `label`
(the leftmost subdomain) instead of a full public hostname; the node's public
domain composes the name. Apex (novox.be) is left as a full name -- composeName
has no empty-label/apex convention yet (mesh-control follow-up).
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Add a `route` contribution (requires/contributes/binds) to every web app so
each gets a Host-routed public name via route-proxy, mirroring the
de-spiegel/only-office pattern:
- novox: gitea, keycloak, nextcloud, umami, invoicing, verdaccio, registry,
and mailu (single mail.novox.be -> 7080; admin/webmail/api ride that port).
- ace: grafana, sonarr, radarr, lidarr, bazarr, ombi, tautulli, jackett,
nodered, searxng, home-assistant, bookshelf, baserow.
Split photos so its three sites each get a name: photos keeps server +
admin-client (photos.novox.be), and new photos-eef (eef.novox.be) and
photos-filip (filip.novox.be) modules carry the client sites.
The production FQDN stays the literal default; a per-node .incus name is a
settings override applied where the mesh runs, not a manifest hardcoding.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The postgres provisioner creates each consumer a database named after the login the mesh
minted (mesh_<node>_<module>), per ADR 0048 -- 'a same-named database under exactly that
login'. But baserow, letta, umami, gitea, keycloak and nextcloud each hardcoded their app db
name (DATABASE_NAME=baserow, /letta, /umami, NAME=gitea, /keycloak, POSTGRES_DB=nextcloud),
so the service connected to a database that does not exist ('database letta does not exist').
Each now uses ${bound:postgres-database:as} as the db name -- the login, which is also the db
name -- matching the working meshboard pattern and the provisioner's actual behaviour.
Found by the two-node DB-consumer lab install (mesh-lab assigned-two-node-db); baserow is
proven connecting and running there. The S3 bucket name is the same class of assumption and is
a separate follow-up (s3 identity has its own length bound).
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Give each settings-config runtime a mergeable config file with its own id
(runtime-config) — the previous "config" collided with modules that already own a
config directory, so ten runtimes mounted a config file no resource declared. Point
each runtime's restart-on at it, so a settings change recreates the runtime and it
re-reads the new value (needs the mesh-host container restart-on fix on
issue/009-container-restart-on). Proven: runtime-restart-on-config e2e green.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Nineteen modules gain a broker-bound runtime container that serves the module's
tools under its own scoped account: bazarr, gitea, grafana, home-assistant,
icecast, influxdb, jackett, keycloak, mailu, nextcloud, nodered, nzbget, ombi,
photos, portainer, qbittorrent, searxng, tautulli, verdaccio.
Config is the assignment's, not the manifest's (ADR 0051): each client's fromEnv
overlays a settings-merged config file (MESH_<M>_CONFIG_FILE) over its env
fallbacks, so URL and credentials come from `settings set`, with the URL defaulting
to the server on the node. nextcloud and mailu also mount the docker socket for
their exec-based tools.
Proven in the mesh-lab: assigned-grafana green — settings deliver the URL and token,
the runtime reads the merged config and serves grafana's tools under the scoped
account, with nothing in the manifest. Two gaps this surfaced are filed as hq
issues 008 (a provider runtime's seal key) and 009 (a settings change does not
restart a container runtime).
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The review found broker own-secret paths drifting: mostly /var/lib/<module>/
broker, but grafana/redis/icecast/nextcloud carried a '-module' suffix to dodge
a collision with the service's own /var/lib/<name> data, and photos sat under
/etc. Normalized to one collision-free namespace a service never owns:
/var/lib/mesh/<module>/broker, with a mesh-state directory resource for the
parent, across 24 modules. audit-logger is grandfathered (lab-proven, referenced
by the assigned test, and it has no service to collide with). All 33 manifests
parse; no client hardcoded a path, so nothing in code moved.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Each module becomes modules/<name>/ holding module.json, with room for
the rest of what a module is — its tools, health checks, provisioning,
lifecycle — which the conversion from hal still has to bring across.
The flat <name>.json was only the resource-declaration half.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF