Implements novox/hq ADR 0109, 0110 and 0111 in the catalogue.
package-registry becomes npm-package-registry throughout (ADR 0109): gitea provides and serves it,
verdaccio provides it, the builder requires, binds and receives its secret under it. gitea's
contributions file is grants/npm.json, so a second ecosystem's file has an obvious name beside it.
gitea claims two mesh seats (ADR 0110): npm-package-registry, which it delivers, and git, which it
now provides with what a clone URL is composed from — http on the forge's web port (ADR 0111).
verdaccio provides npm-package-registry and claims nothing: it is the second provider the seat
exists to make harmless, since a consumer now resolves to the seat's holder without a pin.
No cargo or PyPI provision is added; ADR 0109 defers that. git mints no credential, so gitea's
provisioner registers nothing for it — the mesh's own repositories are public, and a clone
credential is undecided (ADR 0111).
The provisioner still reads where its contributions land from $MESH_RECEIVES, and names no path
itself. One variable carries one path, so a second registration in this module would need the mesh
to say where each provision's file is; that is not possible yet and is not faked here.
Verified: the controller's tests read this catalogue — every claim is a seat in the set, the forge
holds both seats and serves what a clone URL needs, the builder requires what the npm seat delivers
— and pass. Not verified here: a TypeScript build of gitea, whose dependencies resolve from the
private registry.
Two name spaces, two authorities (08-connectivity §2): a public name is
certified by a public CA, an internal one by the mesh's own. step-ca now
offers that second seat as internal-acme-ca beside its existing acme-ca,
and route-proxy requires both — the server dispatches by which authority
may certify the name at all, so an .internal alias stops being plain-HTTP
only without ever asking a public CA for a name it cannot validate.
files-api.novox.be and files.novox.be worked earlier tonight from route-
adapter-generated files, but minio's module.json on main never actually
carried a route requirement — that capability has been sitting in PR #58
the whole time, bundled with an unrelated network rename and console
redirect URL that need their own calmer review. This is just the two
routes: requires: route, contributes.route.api/.console (ContributesMany,
proven working via mesh-controller #55/#57), and the console port (9001)
actually published and declared in listens.
Found assigning route-proxy for the first time tonight: its own routes
file, generated the identical way route-adapter's always was, had four
hostnames in it instead of six — nothing served files-api/files at all,
which would have been a real, silent outage the moment Traefik stopped.
Composed ACME_ROOTS unconditionally from ${bound:acme-ca:roots} even when
that field is empty — public-acme's own case, where an empty roots means
'the system trust store', not 'fetch from the bare authority host'. The
run-once step wget'd https://acme-v02.api.letsencrypt.org:443 (host, no
path) for two minutes every apply and failed, blocking every resource
after it — found live tonight, assigning route-proxy for the first time.
Carries the raw, uncomposed roots value alongside the composed URL
(ACME_ROOTS_PATH) so the step can tell 'nothing to fetch' apart from 'the
authority didn't answer' — a distinction the composed URL alone cannot
make. Empty copies the image's own system CA bundle to /ca/root.crt
instead of fetching one, so ACME_CA_BUNDLE stays the one path it has
always been rather than needing to become conditional itself.
Was still pinned to the scaffold's placeholder digest (mesh-route-
proxy@sha256:0000...0000) even after the build+context work landed —
never caught because nothing had assigned route-proxy before tonight.
artifact: server, matching trust's own reference a few lines up and
every other built module in the catalogue.
step-ca (the other acme-ca provider) already spells it correctly; route-
proxy's own template reads ${bound:acme-ca:roots}. Found live, assigning
public-acme for the first time tonight: the mesh refused the push outright
rather than composing a broken binding — 'route-proxy asks its acme-ca
binding for roots, and what answers it says ... root'. Empty stays empty:
a public CA's root is the system trust store already, per route-proxy's
own design (an empty ACME_CA_BUNDLE means exactly that).
The Dockerfile's own comment already said it — 'the build context is the
mesh-controller repository root' — but nothing in the manifest actually
said so to the mesh, so every build attempt used mesh-catalog's own
directory instead and failed with 'stat go.mod: file does not exist'.
Never caught before because route-proxy has never been assigned anywhere.
Uses the context mechanism just added (mesh-controller#62), proven
working tonight on builder's own self-build.
Its own image ('mesh-builder@sha256:0000...0000', later manually pinned to
a real digest tonight when the placeholder blocked a push) was never
produced by anything the mesh tracks — cmd/mesh-builder lives in
mesh-controller's own repository, and nothing declared how to build an
image from it. Uses the same context mechanism route-proxy does (mesh-
controller#62): the Dockerfile compiles ./cmd/mesh-builder from a clone of
mesh-controller's repository, not a vendored copy.
Unlike mesh-controller's own FROM scratch (ADR 0006 — nothing to audit
but one binary), the build machine's whole job is shelling out to git and
docker, so its runtime is Alpine with both installed from the base's own
packages, not fetched on their own.
Bootstrapped live tonight: a manual build got the new image running long
enough to build itself properly through the pipeline it had just gained,
and mesh-controller itself needed the same upgrade first (it parses
manifests too, and rejected the new context field with the old binary) —
genesis's own kind of ordering problem, solved by hand exactly once.
Bare mesh-builder@sha256:... is only resolvable for a module with a build
section — the mesh's own build step rewrites the reference to a real
registry path as part of resolving build.artifacts. builder is handed
over, not built, so nothing ever rewrites it: pushed as written, docker
read it literally and tried Docker Hub. Took the live node's mesh-builder
down for the length of one push-and-fix (docker: pull access denied for
mesh-builder, repository does not exist). novox.internal:5100, not the
literal external IP docker inspect showed live, for the same reason
addresses generally don't get hardcoded in this catalogue.
The 'package-binding' resource was a hardcoded JSON fragment standing in
for a real grant — {"provision": "package-registry", "from": "gitea",
"at": "127.0.0.1", ...} written as if it were mesh-resolved, when nothing
resolved it. Declares requires: package-registry properly instead, with
binds/secrets pointing at the same file paths the resource used to
manually author, so the mesh mints the grant and writes it there.
npm-password renamed to package-registry.secret: it's gitea's generic
user+password, not npm-specific — the same credential works for basic
auth against cargo/PyPI/Go package endpoints too, once gitea's manifest
grows them (novox/hq ADR 0109).
Known gap, not fixed here (novox/hq issue 117): this makes builder
correct for the steady state but breaks a genesis bootstrap — gitea's own
image is built by builder, so builder cannot yet hold this grant the
first time either has to exist. Filed rather than silently accepted.
2222 never meant anything — no container of gitea's publishes it, the
software never listens on it, and nobody could say where it came from.
The container's real internal sshd is 22 (gitea's own default, unmodified
— every other module's listens.port already means the container's real
internal port, this one didn't). ports now declares 222:22 directly: 222
is the real, fixed, always-known public git-ssh port, so it needs no
per-node setting to reach — unlike an arbitrary auto-assigned port, this
one has external consumers who already know the number.
serves.s3-bucket.region still said us-east-1 (PR #58 already fixed this,
unmerged) while the live mesh-minio sidecar had MESH_MINIO_REGION=eu-west
set out-of-band, not in the manifest at all — and the actual minio server
had no region configured whatsoever (mc admin config get region: empty),
apparently lost across a container recreation since nothing declared it.
Every s3-bucket consumer binding ${bound:s3-bucket:region} was reading
the stale us-east-1 declaration regardless of what was actually live.
Declares MINIO_REGION on the server container and MESH_MINIO_REGION on
the sidecar, matching the static serves declaration, so this is mesh-
managed and durable rather than a manual mc admin config or docker env
override that the next recreation silently drops.
mesh-minio's s3-bucket provisioner shells out to mc to create buckets and
service accounts on the live server, but mc was never in this module's own
runtime image, only in minio's own. It's been silently retrying 'spawn mc
ENOENT' forever, so every s3-bucket grant reached the control-plane layer
(store.json, sealed secret) without the credential ever actually existing
on minio — nextcloud's live instance just hit this as InvalidAccessKeyId
on a real user session.
Copies mc from minio's own image (docker.io/pgsty/minio, already pinned
and pulled as this module's server container) rather than introducing a
new base — mc there is a working, already-verified binary. /usr/bin/mc is
a symlink to mcli; both are copied so it resolves.
the s3-bucket binding already carries serves.region (same mechanism as
at/port); ${bound:s3-bucket:region} tracks whatever minio is actually
configured with instead of a copy that can drift
minio runs with MINIO_REGION=eu-west; nextcloud's S3 config never set a
region, so every object write (avatars, file writes) failed signature
validation with AuthorizationHeaderMalformed, surfacing as Internal Server
Error on real page loads
the migrated data has no literal 'admin' account — HAL's real admin login
is a personal account (jochens), not a generic one. Resetting that would
touch a real user's own credential, so the module gets its own dedicated
admin-group service account instead, same pattern as the minio per-module
service accounts. mesh-admin was created once by hand on novox to match
this manifest for the already-migrated data; a genuinely fresh install
seeds it automatically via NEXTCLOUD_ADMIN_USER/NEXTCLOUD_ADMIN_PASSWORD.
build refused to reach docker:cli implicitly (novox/hq ADR 0097); pin it by
digest and thread it through as DOCKER_CLI, redeclared in the final stage
since args declared before the first FROM don't carry past it
- MESH_NEXTCLOUD_URL hardcoded :80 instead of the mesh-assigned ${port:80}
- sidecar's occ() shells to docker exec but the docker CLI binary was never
present in the runtime image, only the mounted socket
The mesh's own check caught it: an env-file-loaded secret still reaches
the process environment, readable via docker inspect and /proc (hq
04-ISSUES/041) -- the same class of exposure the file-based delivery
exists to avoid. Added MESH_NEXTCLOUD_ADMIN_PASSWORD_FILE support to the
client, matching the pattern the minio client already uses, and mounted
the sealed admin secret directly rather than writing it into an env-file.
The sidecar's own client needs MESH_NEXTCLOUD_ADMIN_PASSWORD to list
shares over the OCS API, but the runtime container's env/volumes never
carried it -- only the server container did. Delivered the same way every
other sealed value in this manifest already is: a generated env-file with
the ${secret:admin} substitution, not a raw value in the container's env.
The pinned digest resolved to a PHP 8.5.10 image; Nextcloud 30 refuses to
run above PHP 8.4. Repinned to the current digest for the nextcloud:30 tag
(matches HAL's own NEXTCLOUD_VERSION), which carries PHP 8.3.28 -- the
same version the data being migrated was actually running under.
minio/minio and minio/mc were pulled from all public registries on 2026-09-11;
pgsty's fork is the working replacement (hq issue 113). The runtime sidecar
had no Dockerfile and no build section at all — added, following the same
tsc-over-client/tools/provisioner shape every other converted module uses.
The data resource pointed straight at /services/minio/data/data1-1, one of
HAL's live 8-drive erasure-coded array — starting this module would have
written into production storage the mesh doesn't own. Moved to a fresh,
empty, mesh-owned directory; the actual data migration happens over the S3
API (rclone), not by sharing a disk path.
hq ADR 0107 / issue 115. HAL's own postgres and lavinmq both used a
directory bind (./db-data, ./data) for exactly this data -- the mesh's
adoption of them, three weeks ago, switched to a named Docker volume
instead, and searxng/distribution followed the same pattern since.
A named volume survives ordinary container recreation, same as a
directory bind -- that was never the problem. The problem is everything
else: docker rm -fv, docker volume rm, and docker system prune --volumes
all target it (one flag away from the docker rm -f this migration already
uses routinely); it is invisible to every tool this migration has used all
night (ls, find, grep across /var/lib, /services); and nothing outside
Docker's own volume machinery can back it up or notice it growing.
mesh-store carries the sharpest version: every database migrated tonight,
including keycloak's, live inside it.
Data already copied and verified on novox before this merges:
- mesh-registry-data -> /var/lib/mesh-registry (11G, diff -rq clean)
- mesh-broker-data -> /var/lib/mesh-broker (37M, cp -a)
- mesh-broker-tls -> /var/lib/mesh-broker-tls (12K, cp -a)
- mesh-store-data -> /var/lib/mesh-store (1.7G) -- mesh-store stopped
cleanly first, so the final copy is crash-consistent, not a live-file
copy of a running postgres; diff -rq clean after.
- searxng/valkey: not yet assigned anywhere, manifest-only fix, nothing
to copy.
Old named volumes left in place, not deleted, as the rollback path.
hq issue 091, measured 2026-09-22: 14 of 46 modules with containers fix the
machine side of a published port in their own manifest -- a fact about one
machine (which port HAL happened to publish it on) written into a
definition meant for any node. ADR 0038 already says the mesh assigns the
machine side and a module says only what it needs; the machinery already
does it (internal/inventory/ports.go's PortFor, declaration.go's
publishedOn rewrites a bare port automatically). These 14 just never
complied.
Fixed 11 of them -- stripped to the bare software port, letting assignment
take over: de-spiegel, gitea, hello-web (the demo module), mailu, mssql,
n8n, novox.be, only-office, photos, photos-eef, photos-filip.
Left alone, the two defensible kinds the issue names: postgres/lavinmq/
distribution (foundation, genesis-rewritten per ADR 0100 -- the number in
the manifest is a default, not a claim) and unifi (protocol/
device-discovery fixes the number; nothing else can find the controller).
gitea was a live one, not just a tidiness fix: its manifest said 2222:22,
but the actual adopted, running container is on 222. Harmless while held,
but the next 'take' would have recreated it on the wrong port and broken
SSH git access. Recorded 222 as novox's own setting for it (settings set
gitea -node novox) so that doesn't happen.
The first commit on this branch embedded ${seat:mesh-store:5432} directly
inside MESH_PROVISION_POSTGRES's URL. That placeholder resolves to an empty
string whenever mesh-store is on its own default port (novox/hq
internal/catalogue/seat_into.go: 'the mesh raised on the catalogue's own
ports never gives them a setting at all') -- which produces a malformed
connection string (host:/postgres) on exactly the common case, a fresh,
non-adopted mesh. It only worked here because novox's mesh-store happens to
be adopted at a non-default port.
mesh-controller's own manifest already has the right shape for this --
internal/envfile/port.go's NAME / NAME_PORT twin, composed by
mesh-controller's Placed(): the base value keeps its own port; a separate
_PORT variable carries the override, spliced in only when it says
something, and left alone -- not a fault -- when it's an unfilled
placeholder (a manifest ahead of the running controller).
Reverted the manifest to its original base value, added
MESH_PROVISION_POSTGRES_PORT as the twin, and taught client.ts (the only
consumer -- both tools/ and provisioner/ import it) the same precedence
envfile.Placed uses. Checked: no other file in the module reads
MESH_PROVISION_POSTGRES directly.
MESH_PROVISION_POSTGRES was a literal connection string naming port 5432 --
correct only when mesh-store happens to run on the mesh's own default. On
novox, mesh-store was adopted in place at HAL's original port (6852), and
the provisioner has been retrying-and-failing against 127.0.0.1:5432 ever
since, for every consumer including ones that already exist (gitea, umami,
mesh-catalog), not just a new grant.
Fixed with the same ${seat:mesh-store:5432} template mesh-controller's own
manifest already uses for the identical connection. No other module needed
this fix checked -- lavinmq's MESH_PROVISION_LAVINMQ already used
127.0.0.1:15672 unconditionally, but the broker's management port is fixed
by the module itself (127.0.0.1:15672 in lavinmq/module.json's own ports),
not by adoption, so it isn't the same bug.
Deployed #49 and the watcher immediately broke: GET /user/repos answered 403,
'required=[read:user]' — confirmed live against the running forge (1.27.3).
That route sits under gitea's user scope category despite listing
repositories, not repository as assumed.
Also gives the fake forge real scope enforcement on /user/repos, which is
why the original PR's test suite didn't catch this: it only checked the
token's value was valid, never that it carried the required scope.
Read against hal/modules/dnsmasq-app (hal dnsmasq-app conversion, hq 08-connectivity).
On the machines it runs, the predecessor's dnsmasq answers every name: the mesh's own
itself, the rest forwarded to 1.1.1.1 and 8.8.8.8, its module's defaults; resolv.conf
names it alone at 127.0.0.1, and the container runtime's dns is the machine's tunnel
address, so the host and every container resolve the world through it. The nox module
forwarded nothing, listened on 127.0.0.55 — a convention of its own beside the one every
machine already followed — and read a machines file that the module's `facts` already
asks the mesh for, so taking it would have left an adopted machine with a resolv.conf
pointing at an address nothing answered on, and no upstream for anything else.
Now the resolver keeps `no-resolv` (the documented loop — finding its own address in
resolv.conf and becoming its own upstream — stays impossible) and forwards to the same
two explicit upstreams; listens on mesh0 and 127.0.0.1, which systemd-resolved does not
hold; requires `mesh-addressing`, since its data is the mesh's addresses; and writes the
runtime's `dns` into daemon.json beside whatever the machine had (ADR 0102), at this
machine's own address — `${machine:address}`, new in the controller. The runtime is not
restarted for it: it reads the key at start, not on reload, and a restart stops every
container; on the machine this replaces the value is already there.
resolv-conf names the resolver alone, as the predecessor's file did; its placeholder second
line was a fallback nothing ever reached. resolved-split-dns follows the address. The
mesh's suffix as a local domain comes with the machines file, so a mesh name the resolver
does not know is refused here rather than asked upstream. mDNS is not carried: no module
does it and the design says mesh names are not multicast names.
The runtime beside the forge served 0 tools: GiteaClient.fromEnv required a token
(settings or MESH_GITEA_TOKEN), nobody had one to give — the mesh raised the forge —
and putting one in settings would store a secret in plaintext in the inventory. So
the fifteen tools registered nothing and the watcher logged "not watching".
What the mesh does deliver is the admin account: a login the manifest names and a
password the vault minted and the host unsealed into a file (ADR 0086). That is
enough to mint a token, so the module does (hq issue 100, the forge's tools):
POST /users/{admin}/tokens over basic auth, scoped to write:repository and
write:issue — the least the tools and the repo watcher need — kept at 0600 in the
module's own state (/var/lib/mesh/gitea/state, a new directory resource the runtime
mounts writable), read back on the next start, and minted afresh when the forge
answers 401 to it or the kept file is gone. A forge whose data came from the
predecessor has no mesh-admin: that is reported in plain words on every poll until
it clears, once per reason, not crash-looped. A configured token still wins and is
never minted over.
The mint happens on the first call, not at registration: a contributor is
synchronous, and a forge not yet answering must not keep the runtime from serving.
One source per kept file in a process, or the watcher and the tools would each
renew on a 401 and drop the other's token by name.
Review of the registry work found it node-scoped: a second `distribution` on another
machine resolved cleanly there, and only afterwards did the mesh notice `artifact-store`
offered by two nodes, with every consumer elsewhere refusing to choose. Worse, a node-scoped
requirement with one candidate installs that candidate, so anything that wanted the store
beside it would have raised a fresh, empty store on the wrong machine first.
There is one store in a mesh, which was already the effective rule; the claim now says it
where a second one is assigned, by name, instead of leaving consumers to discover it.