Every module named its events the way the old bus spelled a routing key —
`module.<module>.<verb>`. Design 29 says a module names an event locally and the
mesh works out where it lands, so all 37 were stale against a rule already
decided. On the new bus that derives into a namespace belonging to a module
called "module", so no cross-module subscription in the mesh matched anything:
nothing failed, nothing reacted (novox/hq 04-ISSUES/127).
36 manifests converted, and 43 files of module code with them. The code mattered
as much as the manifests: the runtime builds the subject from what `emit()` is
handed, so a converted manifest with unconverted code would have had the
permission and the subject disagree.
Three things the new check found on the way:
- `photos` emitted an event its manifest never declared, which the new bus refuses
outright. Declared.
- `showcase` waited for an event nothing emits, so its demo could never be
triggered — only `showcase` may publish under its own name. It emits both halves
now.
- `distribution` declared an event named after a different module. It emits
`image.pushed` under its own name. An event about a *role* belongs on the seat,
where the name outlives whoever holds it, but the sdk has no way to publish on a
seat yet, so that stays recorded rather than declared.
The audit logger's "everything" pattern is `**` rather than the old bus's `#`.
postgres: mesh-store's data directory has always had split ownership --
everything inside pgdata/ is owned by UID 999 (the pgvector image's real
runtime user), while only the top-level mount point happened to be 70:70.
Invisible while the directory's mode was 1777 (world-accessible, from the
named volume this replaced); broke the moment mode: 0700 was enforced,
locking out the actual owning process. mesh-store crash-looped on
Permission denied twice before this was found -- once at container
creation, once mid-session on a checkpoint, after ownership looked correct
by every check that didn't look inside pgdata/ specifically.
keycloak: MESH_KEYCLOAK_URL was hardcoded to :8080, but the module's own
port override (settings set keycloak {ports:{8080:28080}} on novox) means
the real published port is 28080. Same bug class as the postgres
connection-string fix earlier tonight -- now using the mesh's own
template instead, which is exactly the mechanism
internal/catalogue/port_into.go describes for a sidecar dialling its own
server over the machine's loopback.
hq ADR 0107 / issue 115. HAL's own postgres and lavinmq both used a
directory bind (./db-data, ./data) for exactly this data -- the mesh's
adoption of them, three weeks ago, switched to a named Docker volume
instead, and searxng/distribution followed the same pattern since.
A named volume survives ordinary container recreation, same as a
directory bind -- that was never the problem. The problem is everything
else: docker rm -fv, docker volume rm, and docker system prune --volumes
all target it (one flag away from the docker rm -f this migration already
uses routinely); it is invisible to every tool this migration has used all
night (ls, find, grep across /var/lib, /services); and nothing outside
Docker's own volume machinery can back it up or notice it growing.
mesh-store carries the sharpest version: every database migrated tonight,
including keycloak's, live inside it.
Data already copied and verified on novox before this merges:
- mesh-registry-data -> /var/lib/mesh-registry (11G, diff -rq clean)
- mesh-broker-data -> /var/lib/mesh-broker (37M, cp -a)
- mesh-broker-tls -> /var/lib/mesh-broker-tls (12K, cp -a)
- mesh-store-data -> /var/lib/mesh-store (1.7G) -- mesh-store stopped
cleanly first, so the final copy is crash-consistent, not a live-file
copy of a running postgres; diff -rq clean after.
- searxng/valkey: not yet assigned anywhere, manifest-only fix, nothing
to copy.
Old named volumes left in place, not deleted, as the rollback path.
The first commit on this branch embedded ${seat:mesh-store:5432} directly
inside MESH_PROVISION_POSTGRES's URL. That placeholder resolves to an empty
string whenever mesh-store is on its own default port (novox/hq
internal/catalogue/seat_into.go: 'the mesh raised on the catalogue's own
ports never gives them a setting at all') -- which produces a malformed
connection string (host:/postgres) on exactly the common case, a fresh,
non-adopted mesh. It only worked here because novox's mesh-store happens to
be adopted at a non-default port.
mesh-controller's own manifest already has the right shape for this --
internal/envfile/port.go's NAME / NAME_PORT twin, composed by
mesh-controller's Placed(): the base value keeps its own port; a separate
_PORT variable carries the override, spliced in only when it says
something, and left alone -- not a fault -- when it's an unfilled
placeholder (a manifest ahead of the running controller).
Reverted the manifest to its original base value, added
MESH_PROVISION_POSTGRES_PORT as the twin, and taught client.ts (the only
consumer -- both tools/ and provisioner/ import it) the same precedence
envfile.Placed uses. Checked: no other file in the module reads
MESH_PROVISION_POSTGRES directly.
MESH_PROVISION_POSTGRES was a literal connection string naming port 5432 --
correct only when mesh-store happens to run on the mesh's own default. On
novox, mesh-store was adopted in place at HAL's original port (6852), and
the provisioner has been retrying-and-failing against 127.0.0.1:5432 ever
since, for every consumer including ones that already exist (gitea, umami,
mesh-catalog), not just a new grant.
Fixed with the same ${seat:mesh-store:5432} template mesh-controller's own
manifest already uses for the identical connection. No other module needed
this fix checked -- lavinmq's MESH_PROVISION_LAVINMQ already used
127.0.0.1:15672 unconditionally, but the broker's management port is fixed
by the module itself (127.0.0.1:15672 in lavinmq/module.json's own ports),
not by adoption, so it isn't the same bug.
mesh-vault provides `secret` (novox/hq ADR 0085, design 24). The value is
the pair credential the controller mints — the vault holds no copy, only a
ledger of who holds one, its fingerprint and every rotation, and two tools that
answer by fingerprint and never by value. Rotation is `rotate secret`,
unchanged machinery pointed at a secret with an owner (design 13). Named in the
mesh's own namespace, beside mesh-controller and mesh-catalog, because it is
the mesh's own code rather than wrapped software.
redis is the first consumer: its own password stops being an own-secret nothing
could rotate and becomes a `secret` it requires, read from the same file into
the same hole. The server now restarts on its config, or it would keep the
password it started with through every rotation (playbook 06).
keycloak, mailu, minio, mongodb, mssql, nextcloud, portainer, redis,
umami and verdaccio get the Dockerfile + build section the eight
buildable modules already had; their runtime containers name the
artifact instead of a placeholder digest.
One convention, settled (060's open question, informed by 061): the
runtime container runs serve mode with every serve-time entrypoint in
MESH_TOOL_MODULES — tools serve, events flow, and a provider's
provisioner reconciles in the same process with the broker connected.
postgres, gitea and lavinmq are retrofitted from args-run provisioners,
which served no tools and emitted lifecycle events nowhere.
route-proxy is deferred: its build context is the mesh-controller
repository, a cross-repo shape the build section cannot yet express.
postgres claims mesh-store, lavinmq claims mesh-broker, and the controller's seat is
renamed the-controller -> mesh-controller so all three follow one convention. The resolver
refuses a second holder mesh-wide, so an adopted foundation module assigned to a second node
is refused rather than silently raising a second server. Closes hq issue 056 (ADR 0079).
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
server is now the container the foundation raised — same name (mesh-store),
same env (POSTGRES_PASSWORD/PGDATA), ports (127.0.0.1:5432:5432), volume
(mesh-store-data) and the same pinned upstream postgres image the foundation
runs — so the applier adopts it in place rather than raising a second postgres.
The provisioner is host-networked to reach the loopback store at 127.0.0.1:5432.
The module's own network and bind-mounted data dir are gone; there is one
postgres now, holding the controller's contexts and every module's database.
Issue 051 (WBS 3.1).
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Each named a digest produced inside a lab that no longer exists, so none of them
could be built anywhere else. They say which module they stand on now, and there
is deliberately no default — a build nobody told stops at the declaration rather
than at a reference that resolves to nothing.
Its provisioner container named an image nobody could produce — a zero digest
placeholder. It names an artifact instead, and the module says how to build it,
so the mesh can make the database provider the catalogue needs.
The breadth install (catalogue-broad) surfaced two latent resolution bugs that
single-module compile checks never caught (compile != resolve):
1. postgres/redis/mongodb/minio served no "port", yet seven consumers
(baserow, letta, invoicing, gitea, umami, keycloak, nextcloud) reference
${bound:<provision>:port}. A binding auto-carries at/from/as; the port is the
provider's half of the answer and must be declared in `serves` (manifest.go:
"Serves is what a consumer needs to know ... a port, a path, a realm"). Added
port to each: postgres 5432, redis 6379, mongodb 27017, minio 9000. Without it
no DB consumer could resolve, let alone deploy.
2. baserow declared an empty contribution `contributes: {"redis-cache": {}}`.
redis's serves spec for the cache is empty (a cache takes no per-consumer
payload), so a consumer only `requires` it; an empty contribution is refused.
Dropped it (requires/binds/secrets unchanged).
Found by mesh-lab catalogue-broad; the same fixes are mirrored in that bed.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Each provider's separate provisioner container becomes a runtime container that
serves the module's tools and runs its provisioner under the module's scoped broker
account: mesh-runtime-<module>, on the backend's own network (reaching the backend
by name and the broker by NAT), with MESH_BROKER_FILE + MESH_RECEIVES replacing
GRANTS. umami gains the broker own-secret it lacked. cloudflare-dns's adapter is
re-pointed at the ADR 0053 contract (a data provision, like umami — its record
return is the scoped-out concern).
Proven: provider-on-backend-network green — redis's runtime, on the private redis
network, binds the broker and provisions a consumer with the mesh's credential.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
plex gains a broker-bound tools/events runtime container (mesh-runtime-plex)
alongside its server, and a never-throwing plex_reachable health probe. redis and
postgres subscribe to their own lifecycle events in index.ts but declared no
consumes — so the substrate never made the queue the runtime binds and it crashed
on start (404). Declare the consume, as ADR 0046 requires. Proven end-to-end in the
mesh-lab: assigned-plex and assigned-redis both green.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The review found broker own-secret paths drifting: mostly /var/lib/<module>/
broker, but grafana/redis/icecast/nextcloud carried a '-module' suffix to dodge
a collision with the service's own /var/lib/<name> data, and photos sat under
/etc. Normalized to one collision-free namespace a service never owns:
/var/lib/mesh/<module>/broker, with a mesh-state directory resource for the
parent, across 24 modules. audit-logger is grandfathered (lab-proven, referenced
by the assigned test, and it has no service to collide with). All 33 manifests
parse; no client hardcoded a path, so nothing in code moved.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
redis provides redis-cache: a real admin client speaking RESP over a raw socket
(node:net, no deps); provisioner makes a keyspace-scoped ACL user per grant.
postgres provides postgres-database: admin client executing through psql (the
consistent shell-out port, like minio's mc), full DDL for create/drop database+
role, a CSV row parser, read-only query tool. Both emit
module.<x>.<thing>.provisioned/.deprovisioned from the provisioner. Both
typecheck; manifests parse.
Each module becomes modules/<name>/ holding module.json, with room for
the rest of what a module is — its tools, health checks, provisioning,
lifecycle — which the conversion from hal still has to bring across.
The flat <name>.json was only the resource-declaration half.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF