The env block carried the key twice — mailu-webmail from #79's address
sweep, and a stray =webmail further down that survived it. Last write
wins in an env file, so the front resolved a name that answers nowhere
on the mesh's network and 502'd every logged-in webmail request. Latent
since the cutover: the SSO redirect the checks watched never touches
the upstream; the operator's first real login did.
Take one (#90) died on two real edge bugs, both fixed and pinned by
tests in mesh-controller (#66: autocert 404s unknown tokens itself;
#67: the internal authority 403s every public name before the token
lookup). The challenge path verified end to end reaching mailu's own
nginx before this flip.
The letsencrypt flavor served certbot's April-expired state to live IMAPS
users within minutes: autocert's HTTPHandler answers 404 itself for
tokens it does not hold and never consults the fallback for challenge
paths, so mailu's own client cannot answer through the path-scoped
route. cert flavor (valid to Nov 27) until route-proxy's handler
actually falls through.
PR #82 set TLS_FLAVOR=cert as the honest interim while the predecessor's
proxy owned /.well-known/acme-challenge outright. route-proxy took port
80 today and its handler passes unknown tokens through to routed paths
by design — the one line #82 promised, made now. The copied cert (valid
to Nov 27) stays on disk untouched; mailu's own certbot takes over from
here.
The manifest predated the working deployment on three axes: it declared a
data directory the running portainer never used (taking it would have
started empty), pinned an image digest the machine has moved past (issue
099), and contributed no route while portainer.novox.be rides a traefik
container label today. Now: the predecessor's portainer_data path, the
running image's digest, 9090:9000 kept as the predecessor's machine port
with the route contribution naming it, and 9443 kept for the runtime
sidecar's own TLS conversation.
A bare '80' tried to bind the node's port 80 — the edge's — instead of
auto-allocating. 9070 is the predecessor's number and the one the route
contribution already names.
The app reads MONGO_DB (default 'invoicing') for every operation and
uses the URL only to connect — listCollections ran against a database
the granted user cannot see. Same fault and same fix as photos' MONGO_DB,
found by the API's own logs at take.
The mongo credential authenticates against its own database and the
database is the granted one (mesh_novox_invoice), not the contributed
name the provisioner ignores. Same for the store: the key is sealed to
the derived bucket (mesh-novox-invoice) — the data mirrors in during the
window, the ncloud/photos pattern. And the api gets the route
contribution it always needed: invoicing-api.novox.be is today a traefik
container label, invisible to every file survey, and it must be a grant
before the edge can ever flip.
The adopted node still runs the predecessor's mongo container, and it must
keep running: invoicing points at novox.be:27017 and is not migrating in
this window. A module container named 'mongo' would be held at assign and
would replace the predecessor at take, cutting invoicing off its database.
The mesh's server coexists instead — fresh data directory, its own name,
auto-allocated machine port — and the predecessor retires with its last
consumer.
The declared 0700 was applied at take and broke mail quietly: postfix's
master runs as root but pickup and smtpd drop to uid postfix, and a
spool root they cannot traverse is a maildrop they cannot scan and a
rewrite socket they cannot open — auth succeeded and MAIL FROM hung.
0755 root is exactly what postfix's own set-permissions makes of
/var/spool/postfix. Fixed live by chmod first; declared here so the
next push stops undoing it.
letsencrypt was the aspiration and cannot work yet, proven live: the
predecessor's own ACME machinery owns /.well-known/acme-challenge on
port 80 outright (unknown tokens get its 404) and its entrypoint
redirect owns every other path — the hand-authored passthrough never
matched anything, which is why mailu's certbot state had quietly
expired in April while the copied files carried the name. cert flavor
serves those files (valid to Nov 27). Mailu certifying itself becomes
possible the day route-proxy takes port 80, whose handler falls through
unknown tokens by design — that flip is one line here, made then.
The hand-written seed predates automx2's prio column, so every
config-v1.1.xml request 500'd against a table the seed had just made —
and the predecessor's own database had the same gap: client
autoconfiguration has been silently broken on the old stack for a long
time, behind a root page that answered 200. The live database gained
the column by ALTER; a fresh mesh now seeds it right.
An unpinned pip install took the latest automx2, whose schema grew a
column (server.prio) the module's own seeding SQL predates — a 500 on
every autoconfig request against a table the seed had just written.
2021.6 is what the proven image runs; the seed and the software agree
again. Regenerating the seed for a newer automx2 is its own change,
made deliberately, not by whatever pip resolved this week.
The front resolves its upstreams from *_ADDRESS, defaulting to the bare
compose service names — admin, antispam — and reads the 1.9-era HOST_*
not at all. Phase 1 never noticed because the predecessor's service
names WERE the defaults; the mesh's containers are mailu-*, and the
front answered 502 asking docker for a name nothing carries. The dead
vocabulary goes; every upstream is named as the container actually is.
2024.06's admin refuses to serve behind a resolver that does not
validate DNSSEC — found live as an unhealthy admin, 454s on submission
and a 500 webmail, with the runtime's forwarder validating nothing. The
unbound this module always shipped becomes reachable: pinned at the
predecessor's own address on the module network (the subnet the mesh
adopted), and named as dns by the nine containers that resolve anything.
Stands on mesh-host #26, which gave the vocabulary these two fields.
1.9's admin served on 80; 2024.06's gunicorn listens on 8080, and the
runtime's fetch failed against the old port the moment the broker
credential let it try.
The predecessor's image carried .venv/scripts/flask.sh, created by a
build step that never made it into the files this module holds — the
image worked and its recipe could not reproduce it, caught the moment
the mesh built it from source (exit 127 crash loop at cutover). The
wrapper is now written explicitly, verbatim from the proven image, so
the recipe is the whole truth about the image.
Phase 1 (MAILU-CUTOVER.md) upgraded the live stack 1.9→2024.06 and its
lessons land here as pins: every image is the digest running and
verified tonight — webmail under the name 2024.06 actually uses (the
draft's roundcube pin misled a whole hop), antivirus on the upstream
clamav image with the signature DB in its own directory (the mailu-built
image ended at 2.0), and the front trusting the proxy address the live
config actually names. The welcome-mail texts ride along for parity.
TLS_FLAVOR stays letsencrypt deliberately where the live .env says cert:
the cert files are copies whose HAL-era renewal hook died with HAL
(expiry Nov 27); mailu managing its own issuance through the existing
ACME passthrough is the fix, and the files remain on disk as the
fallback flavor if first issuance misbehaves during the window.
PORTS defaults to 25,80,443,465,993,995,4190 in 2024.06 — submission on
587 among the closed, which is what every client of this server uses.
The same parity decision the listens already state, now stated where the
software reads it.
Five gaps between the draft and what actually runs, each verified live
before being written down:
- front published bare 80 — the machine port Traefik holds; now the
predecessor's own mappings (7080:80, 7443:443) plus the 110/143/995
parity ports the draft dropped. Pruning legacy protocols is its own
deliberate change, not a cutover side effect.
- TLS_FLAVOR said cert, which nothing supplies; live is letsencrypt —
mailu runs its own certbot, state already on disk, HTTP-01 answered
through a path-scoped route contribution (priority above the web one).
- the web route said http:7080, the redirect-loop shape; it now says
what the hand-authored file always knew: https 7443, insecure.
- automx was absent entirely: the autoconfig responder is now a second
artifact (its Containerfile moved in from the predecessor's images
dir, base declared per ADR 0097), a container on a real data dir —
the anonymous-volume loss of 2026-08-10 stays fixed — and the three
public names are route contributions.
- and the reason this moved ahead of de-spiegel: mailu now provides
smtp. A consumer contributes the account it sends as; the provisioner
creates <account>@<domain> via the admin API and applies the minted
password every reconcile (ADR 0048). The domain is served on the
binding so a consumer composes its own login from mesh facts.
route-adapter learns to say no: a contribution over https, scoped to a
path, or carrying a policy is skipped aloud rather than written into a
file shape that cannot say it — plain http into a TLS listener was the
concrete wrong file this prevents. The hand-authored files keep covering
those routes until the mesh's own proxy takes over, exactly as today.
The builder's first credentialed clone of a private repository answered
'not found': the packages team named only repo.packages in its
units_map, which is exhaustive — so members had no code unit at all, and
gitea hides what a user cannot read. One credential answering npm and
git alike was the whole design of the builder's grant; the team now says
so.
And found teams are patched, not just returned: a team is configuration
the reconcile loop owns, the same as a user's password, so a unit this
code gains reaches the team that already exists rather than only the
next mesh raised from scratch.
The builder's package-registry grant — the first this provider ever
received — retried for a day saying only that an edit 404'd. Two faults
under it: the API refuses an email without a dotted domain, so
`@localhost` failed validation at create (the CLI that made mesh-admin
accepts it, which is why the admin exists and no consumer did); and
ensureUser read that 422 as 'already exists' and went on to edit a user
that was never made, burying the create's own message. The address is now
gitea's own hidden-address shape, and the edit path is taken only for a
user that is actually there.
The 2026-09-12 incident response blocked /api/internal by hand in the
predecessor's dynamic directory, with a note that its durable home is the
mesh's routing. A route carries the policy applied to a request (ADR
0108), so the refusal now travels with the grant: route-proxy enforces
it on both the public name and the internal alias the moment it serves
this route, and the adapter skips it aloud (no port, nothing to write)
while the predecessor's own file still stands. The hand-authored file
retires with the proxy it configures.
Two name spaces, two authorities (08-connectivity §2): a public name is
certified by a public CA, an internal one by the mesh's own. step-ca now
offers that second seat as internal-acme-ca beside its existing acme-ca,
and route-proxy requires both — the server dispatches by which authority
may certify the name at all, so an .internal alias stops being plain-HTTP
only without ever asking a public CA for a name it cannot validate.
files-api.novox.be and files.novox.be worked earlier tonight from route-
adapter-generated files, but minio's module.json on main never actually
carried a route requirement — that capability has been sitting in PR #58
the whole time, bundled with an unrelated network rename and console
redirect URL that need their own calmer review. This is just the two
routes: requires: route, contributes.route.api/.console (ContributesMany,
proven working via mesh-controller #55/#57), and the console port (9001)
actually published and declared in listens.
Found assigning route-proxy for the first time tonight: its own routes
file, generated the identical way route-adapter's always was, had four
hostnames in it instead of six — nothing served files-api/files at all,
which would have been a real, silent outage the moment Traefik stopped.
Composed ACME_ROOTS unconditionally from ${bound:acme-ca:roots} even when
that field is empty — public-acme's own case, where an empty roots means
'the system trust store', not 'fetch from the bare authority host'. The
run-once step wget'd https://acme-v02.api.letsencrypt.org:443 (host, no
path) for two minutes every apply and failed, blocking every resource
after it — found live tonight, assigning route-proxy for the first time.
Carries the raw, uncomposed roots value alongside the composed URL
(ACME_ROOTS_PATH) so the step can tell 'nothing to fetch' apart from 'the
authority didn't answer' — a distinction the composed URL alone cannot
make. Empty copies the image's own system CA bundle to /ca/root.crt
instead of fetching one, so ACME_CA_BUNDLE stays the one path it has
always been rather than needing to become conditional itself.
Was still pinned to the scaffold's placeholder digest (mesh-route-
proxy@sha256:0000...0000) even after the build+context work landed —
never caught because nothing had assigned route-proxy before tonight.
artifact: server, matching trust's own reference a few lines up and
every other built module in the catalogue.
step-ca (the other acme-ca provider) already spells it correctly; route-
proxy's own template reads ${bound:acme-ca:roots}. Found live, assigning
public-acme for the first time tonight: the mesh refused the push outright
rather than composing a broken binding — 'route-proxy asks its acme-ca
binding for roots, and what answers it says ... root'. Empty stays empty:
a public CA's root is the system trust store already, per route-proxy's
own design (an empty ACME_CA_BUNDLE means exactly that).
The Dockerfile's own comment already said it — 'the build context is the
mesh-controller repository root' — but nothing in the manifest actually
said so to the mesh, so every build attempt used mesh-catalog's own
directory instead and failed with 'stat go.mod: file does not exist'.
Never caught before because route-proxy has never been assigned anywhere.
Uses the context mechanism just added (mesh-controller#62), proven
working tonight on builder's own self-build.
Its own image ('mesh-builder@sha256:0000...0000', later manually pinned to
a real digest tonight when the placeholder blocked a push) was never
produced by anything the mesh tracks — cmd/mesh-builder lives in
mesh-controller's own repository, and nothing declared how to build an
image from it. Uses the same context mechanism route-proxy does (mesh-
controller#62): the Dockerfile compiles ./cmd/mesh-builder from a clone of
mesh-controller's repository, not a vendored copy.
Unlike mesh-controller's own FROM scratch (ADR 0006 — nothing to audit
but one binary), the build machine's whole job is shelling out to git and
docker, so its runtime is Alpine with both installed from the base's own
packages, not fetched on their own.
Bootstrapped live tonight: a manual build got the new image running long
enough to build itself properly through the pipeline it had just gained,
and mesh-controller itself needed the same upgrade first (it parses
manifests too, and rejected the new context field with the old binary) —
genesis's own kind of ordering problem, solved by hand exactly once.
Bare mesh-builder@sha256:... is only resolvable for a module with a build
section — the mesh's own build step rewrites the reference to a real
registry path as part of resolving build.artifacts. builder is handed
over, not built, so nothing ever rewrites it: pushed as written, docker
read it literally and tried Docker Hub. Took the live node's mesh-builder
down for the length of one push-and-fix (docker: pull access denied for
mesh-builder, repository does not exist). novox.internal:5100, not the
literal external IP docker inspect showed live, for the same reason
addresses generally don't get hardcoded in this catalogue.
The 'package-binding' resource was a hardcoded JSON fragment standing in
for a real grant — {"provision": "package-registry", "from": "gitea",
"at": "127.0.0.1", ...} written as if it were mesh-resolved, when nothing
resolved it. Declares requires: package-registry properly instead, with
binds/secrets pointing at the same file paths the resource used to
manually author, so the mesh mints the grant and writes it there.
npm-password renamed to package-registry.secret: it's gitea's generic
user+password, not npm-specific — the same credential works for basic
auth against cargo/PyPI/Go package endpoints too, once gitea's manifest
grows them (novox/hq ADR 0109).
Known gap, not fixed here (novox/hq issue 117): this makes builder
correct for the steady state but breaks a genesis bootstrap — gitea's own
image is built by builder, so builder cannot yet hold this grant the
first time either has to exist. Filed rather than silently accepted.
2222 never meant anything — no container of gitea's publishes it, the
software never listens on it, and nobody could say where it came from.
The container's real internal sshd is 22 (gitea's own default, unmodified
— every other module's listens.port already means the container's real
internal port, this one didn't). ports now declares 222:22 directly: 222
is the real, fixed, always-known public git-ssh port, so it needs no
per-node setting to reach — unlike an arbitrary auto-assigned port, this
one has external consumers who already know the number.
serves.s3-bucket.region still said us-east-1 (PR #58 already fixed this,
unmerged) while the live mesh-minio sidecar had MESH_MINIO_REGION=eu-west
set out-of-band, not in the manifest at all — and the actual minio server
had no region configured whatsoever (mc admin config get region: empty),
apparently lost across a container recreation since nothing declared it.
Every s3-bucket consumer binding ${bound:s3-bucket:region} was reading
the stale us-east-1 declaration regardless of what was actually live.
Declares MINIO_REGION on the server container and MESH_MINIO_REGION on
the sidecar, matching the static serves declaration, so this is mesh-
managed and durable rather than a manual mc admin config or docker env
override that the next recreation silently drops.
OBJECTSTORE_S3_BUCKET=nextcloud was a leftover from before the module
existed — that bucket was created by hand during tonight's earlier HAL
credential stopgap. The mesh's own minio provisioner derives its own
bucket name from the consumer's access-key identity (bucketFor(as) in
minio/client.ts) rather than honouring contributes.s3-bucket.bucket — by
design, so teardown can recompute the name with nothing persisted — and
minted mesh-novox-ncloud, a different bucket. mesh_novox_ncloud's scoped
policy only covers that bucket, so every S3 write 403'd with AccessDenied
trying to touch the old one. Pointed both the request hint and the real
env var at the bucket that's actually there.
mesh-minio's s3-bucket provisioner shells out to mc to create buckets and
service accounts on the live server, but mc was never in this module's own
runtime image, only in minio's own. It's been silently retrying 'spawn mc
ENOENT' forever, so every s3-bucket grant reached the control-plane layer
(store.json, sealed secret) without the credential ever actually existing
on minio — nextcloud's live instance just hit this as InvalidAccessKeyId
on a real user session.
Copies mc from minio's own image (docker.io/pgsty/minio, already pinned
and pulled as this module's server container) rather than introducing a
new base — mc there is a working, already-verified binary. /usr/bin/mc is
a symlink to mcli; both are copied so it resolves.
the s3-bucket binding already carries serves.region (same mechanism as
at/port); ${bound:s3-bucket:region} tracks whatever minio is actually
configured with instead of a copy that can drift
minio runs with MINIO_REGION=eu-west; nextcloud's S3 config never set a
region, so every object write (avatars, file writes) failed signature
validation with AuthorizationHeaderMalformed, surfacing as Internal Server
Error on real page loads
the migrated data has no literal 'admin' account — HAL's real admin login
is a personal account (jochens), not a generic one. Resetting that would
touch a real user's own credential, so the module gets its own dedicated
admin-group service account instead, same pattern as the minio per-module
service accounts. mesh-admin was created once by hand on novox to match
this manifest for the already-migrated data; a genuinely fresh install
seeds it automatically via NEXTCLOUD_ADMIN_USER/NEXTCLOUD_ADMIN_PASSWORD.
build refused to reach docker:cli implicitly (novox/hq ADR 0097); pin it by
digest and thread it through as DOCKER_CLI, redeclared in the final stage
since args declared before the first FROM don't carry past it
- MESH_NEXTCLOUD_URL hardcoded :80 instead of the mesh-assigned ${port:80}
- sidecar's occ() shells to docker exec but the docker CLI binary was never
present in the runtime image, only the mounted socket
The mesh's own check caught it: an env-file-loaded secret still reaches
the process environment, readable via docker inspect and /proc (hq
04-ISSUES/041) -- the same class of exposure the file-based delivery
exists to avoid. Added MESH_NEXTCLOUD_ADMIN_PASSWORD_FILE support to the
client, matching the pattern the minio client already uses, and mounted
the sealed admin secret directly rather than writing it into an env-file.
The sidecar's own client needs MESH_NEXTCLOUD_ADMIN_PASSWORD to list
shares over the OCS API, but the runtime container's env/volumes never
carried it -- only the server container did. Delivered the same way every
other sealed value in this manifest already is: a generated env-file with
the ${secret:admin} substitution, not a raw value in the container's env.
The pinned digest resolved to a PHP 8.5.10 image; Nextcloud 30 refuses to
run above PHP 8.4. Repinned to the current digest for the nextcloud:30 tag
(matches HAL's own NEXTCLOUD_VERSION), which carries PHP 8.3.28 -- the
same version the data being migrated was actually running under.
Left at MinIO's us-east-1 default. Novox is hosted in Germany, the team
is in Belgium -- eu-west is correct, and matters beyond labeling: it's
part of the SigV4 signature, so a client using the wrong region fails
auth even with valid credentials. Set on the server (MINIO_REGION),
the served provision value, and the runtime sidecar's own client.
The 4-node/8-drive erasure-coded cluster matched HAL's topology faithfully,
but real throughput testing against both showed why that costs more than
it's worth here: every write on the sharded cluster fans out across 4
processes over the internal network with erasure-coding overhead, capping
safe throughput around 1.3-2 MiB/s and breaking outright above ~256
concurrent transfers (IncompleteBody errors, confirmed via a controlled
512x test). The identical copy against a single-node instance sustained
23+ MiB/s at the same concurrency with zero errors — over 10x faster,
verified side-by-side, not assumed.
Trades away erasure-coded redundancy (no single-drive fault tolerance) for
that throughput. Deliberate, and reversible if it turns out to matter later
-- the data itself is migrated over the S3 API either way, so the storage
topology underneath isn't locked in by anything upstream of it.
Collided with the LB container's own name. docker inspect minio resolved
to the network instead of the (not-yet-created) container, and mesh-host's
existence check crashed on the mismatched shape rather than reporting
absence -- a real mesh-host bug (fixed separately, mesh-host#25), but this
sidesteps it here without waiting on a host-level binary update.
files-api.novox.be (port 9000, the S3 data API) and files.novox.be (port
9001, the console) — same two names HAL routes today, via nginx's own
upstream split. Needed mesh-controller#55 (a module answering one
requirement several times) to exist first; it's merged and deployed.
The single standalone instance from the first pass didn't match HAL's actual
topology: HAL runs minio1-4, two drives each, behind an nginx load balancer
on 9000 (S3) and 9001 (console). This rewrite mirrors that exactly — same
node count, same erasure-coding command, same LB config — so the migration
is a real like-for-like move, not a simplification.
Only the two images that had to change did: the minio server (dead upstream,
already fixed in the prior commit) and nginx (1.19.2-alpine is long EOL;
repinned to current stable-alpine by digest). Data still lands on a fresh,
empty, mesh-owned path, never HAL's live drives.
The OIDC-wait entrypoint wrapper HAL used is dropped: it's a no-op when
MINIO_IDENTITY_OPENID_CONFIG_URL is unset (it always is here — no OIDC
integration was ever wired to minio itself), and this catalogue has no
container resource field for overriding a container's entrypoint anyway —
every converted module relies on the image's own entrypoint plus args,
which is exactly what the original single-node version already did.
minio/minio and minio/mc were pulled from all public registries on 2026-09-11;
pgsty's fork is the working replacement (hq issue 113). The runtime sidecar
had no Dockerfile and no build section at all — added, following the same
tsc-over-client/tools/provisioner shape every other converted module uses.
The data resource pointed straight at /services/minio/data/data1-1, one of
HAL's live 8-drive erasure-coded array — starting this module would have
written into production storage the mesh doesn't own. Moved to a fresh,
empty, mesh-owned directory; the actual data migration happens over the S3
API (rclone), not by sharing a disk path.
The previous commit on this branch used KC_PROXY=edge and
KC_HOSTNAME_STRICT_HTTPS=true, carried over from HAL's config -- but
HAL ran an older Keycloak using the v1 hostname provider. This image
(26.0.8) defaults to Hostname v2, which warned 'options [proxy,
hostname-strict-https] are still in use, please review your
configuration' and kept generating http:// URLs regardless -- verified
against /realms/Novox/.well-known/openid-configuration directly, not
just the login button, after the first fix deployed.
v2's actual shape (keycloak.org/server/hostname): KC_HOSTNAME is a full
URL, not a bare hostname -- the scheme in the URL is what tells Keycloak
to generate https, not a separate strict-https flag. KC_PROXY_HEADERS
replaces KC_PROXY: xforwarded to trust traefik's X-Forwarded-* headers,
which it sends by default.
Reported: files.novox.be's login button redirects to http://keycloak.novox.be,
not https. HAL's original config (/services/keycloak/docker-compose.yml) set
three settings the mesh's manifest never carried over:
KC_HOSTNAME: keycloak.novox.be
KC_HOSTNAME_STRICT_HTTPS: true
KC_PROXY: edge
Without KC_PROXY: edge, Keycloak has no way to know it sits behind a
TLS-terminating reverse proxy (traefik) -- it generates URLs from what it
directly sees, which is plain HTTP from traefik's backend connection. Same
pattern as the named-volume conversion: the shape was rebuilt from general
knowledge of what a keycloak container needs, not from what this
installation's own working config actually had.
postgres: mesh-store's data directory has always had split ownership --
everything inside pgdata/ is owned by UID 999 (the pgvector image's real
runtime user), while only the top-level mount point happened to be 70:70.
Invisible while the directory's mode was 1777 (world-accessible, from the
named volume this replaced); broke the moment mode: 0700 was enforced,
locking out the actual owning process. mesh-store crash-looped on
Permission denied twice before this was found -- once at container
creation, once mid-session on a checkpoint, after ownership looked correct
by every check that didn't look inside pgdata/ specifically.
keycloak: MESH_KEYCLOAK_URL was hardcoded to :8080, but the module's own
port override (settings set keycloak {ports:{8080:28080}} on novox) means
the real published port is 28080. Same bug class as the postgres
connection-string fix earlier tonight -- now using the mesh's own
template instead, which is exactly the mechanism
internal/catalogue/port_into.go describes for a sidecar dialling its own
server over the machine's loopback.
hq ADR 0107 / issue 115. HAL's own postgres and lavinmq both used a
directory bind (./db-data, ./data) for exactly this data -- the mesh's
adoption of them, three weeks ago, switched to a named Docker volume
instead, and searxng/distribution followed the same pattern since.
A named volume survives ordinary container recreation, same as a
directory bind -- that was never the problem. The problem is everything
else: docker rm -fv, docker volume rm, and docker system prune --volumes
all target it (one flag away from the docker rm -f this migration already
uses routinely); it is invisible to every tool this migration has used all
night (ls, find, grep across /var/lib, /services); and nothing outside
Docker's own volume machinery can back it up or notice it growing.
mesh-store carries the sharpest version: every database migrated tonight,
including keycloak's, live inside it.
Data already copied and verified on novox before this merges:
- mesh-registry-data -> /var/lib/mesh-registry (11G, diff -rq clean)
- mesh-broker-data -> /var/lib/mesh-broker (37M, cp -a)
- mesh-broker-tls -> /var/lib/mesh-broker-tls (12K, cp -a)
- mesh-store-data -> /var/lib/mesh-store (1.7G) -- mesh-store stopped
cleanly first, so the final copy is crash-consistent, not a live-file
copy of a running postgres; diff -rq clean after.
- searxng/valkey: not yet assigned anywhere, manifest-only fix, nothing
to copy.
Old named volumes left in place, not deleted, as the rollback path.
hq issue 091, measured 2026-09-22: 14 of 46 modules with containers fix the
machine side of a published port in their own manifest -- a fact about one
machine (which port HAL happened to publish it on) written into a
definition meant for any node. ADR 0038 already says the mesh assigns the
machine side and a module says only what it needs; the machinery already
does it (internal/inventory/ports.go's PortFor, declaration.go's
publishedOn rewrites a bare port automatically). These 14 just never
complied.
Fixed 11 of them -- stripped to the bare software port, letting assignment
take over: de-spiegel, gitea, hello-web (the demo module), mailu, mssql,
n8n, novox.be, only-office, photos, photos-eef, photos-filip.
Left alone, the two defensible kinds the issue names: postgres/lavinmq/
distribution (foundation, genesis-rewritten per ADR 0100 -- the number in
the manifest is a default, not a claim) and unifi (protocol/
device-discovery fixes the number; nothing else can find the controller).
gitea was a live one, not just a tidiness fix: its manifest said 2222:22,
but the actual adopted, running container is on 222. Harmless while held,
but the next 'take' would have recreated it on the wrong port and broken
SSH git access. Recorded 222 as novox's own setting for it (settings set
gitea -node novox) so that doesn't happen.