hq ADR 0107 / issue 115. HAL's own postgres and lavinmq both used a directory bind (./db-data, ./data) for exactly this data — the mesh's adoption of them three weeks ago switched to a named Docker volume instead, and distribution/searxng followed the same pattern since.
Already live on novox, not just proposed — this was urgent enough (operator instruction, mid-session) to build and roll out before opening this PR. Full account of what happened:
Data copied and verified before any manifest change applied: mesh-registry-data → /var/lib/mesh-registry (11G, diff -rq clean), mesh-broker-data/mesh-broker-tls → /var/lib/mesh-broker(-tls) (37M + 12K), mesh-store-data → /var/lib/mesh-store (1.7G) — mesh-store stopped cleanly first so the final copy is crash-consistent, not a live-file copy of a running postgres; diff -rq clean after.
Incident during rollout, self-inflicted and resolved: the new host directories were created via sudo, so mesh-store crash-looped on mkdir: cannot create directory '/var/lib/postgresql/data': Permission denied — a host bind mount doesn't get auto-chowned by Docker the way a named volume does. Fixed by matching the original volume's ownership (70:70, mode 1777) onto the new directory. mesh-broker and mesh-registry didn't hit this — their copied files' ownership was already compatible. Total mesh-store downtime: a few minutes, during which mesh-controller itself couldn't reach its own inventory (expected — it's also on mesh-store) and briefly restart-looped; recovered on its own once mesh-store came back.
Verified after recovery: all four post-migration databases present and correct (mesh_novox_gitea, mesh_novox_umami, mesh_novox_catalog, mesh_novox_keycloak), keycloak's realm/user counts unchanged (2 realms, 13 users, matching the count from the restore earlier tonight), mesh-broker healthy with active connections, mesh-registry serving, gitea's sidecar self-healed.
Old named volumes deliberately left in place, not deleted — the rollback path, per standing instruction to always be able to roll back.
hq ADR 0107 / issue 115. HAL's own `postgres` and `lavinmq` both used a directory bind (`./db-data`, `./data`) for exactly this data — the mesh's adoption of them three weeks ago switched to a named Docker volume instead, and `distribution`/`searxng` followed the same pattern since.
**Already live on `novox`, not just proposed** — this was urgent enough (operator instruction, mid-session) to build and roll out before opening this PR. Full account of what happened:
- Data copied and verified before any manifest change applied: `mesh-registry-data` → `/var/lib/mesh-registry` (11G, `diff -rq` clean), `mesh-broker-data`/`mesh-broker-tls` → `/var/lib/mesh-broker(-tls)` (37M + 12K), `mesh-store-data` → `/var/lib/mesh-store` (1.7G) — `mesh-store` stopped cleanly first so the final copy is crash-consistent, not a live-file copy of a running postgres; `diff -rq` clean after.
- **Incident during rollout, self-inflicted and resolved:** the new host directories were created via `sudo`, so `mesh-store` crash-looped on `mkdir: cannot create directory '/var/lib/postgresql/data': Permission denied` — a host bind mount doesn't get auto-chowned by Docker the way a named volume does. Fixed by matching the original volume's ownership (`70:70`, mode `1777`) onto the new directory. `mesh-broker` and `mesh-registry` didn't hit this — their copied files' ownership was already compatible. Total `mesh-store` downtime: a few minutes, during which `mesh-controller` itself couldn't reach its own inventory (expected — it's also on `mesh-store`) and briefly restart-looped; recovered on its own once `mesh-store` came back.
- Verified after recovery: all four post-migration databases present and correct (`mesh_novox_gitea`, `mesh_novox_umami`, `mesh_novox_catalog`, `mesh_novox_keycloak`), `keycloak`'s realm/user counts unchanged (2 realms, 13 users, matching the count from the restore earlier tonight), `mesh-broker` healthy with active connections, `mesh-registry` serving, `gitea`'s sidecar self-healed.
- Old named volumes deliberately left in place, not deleted — the rollback path, per standing instruction to always be able to roll back.
`searxng`/`valkey` isn't assigned anywhere yet — manifest-only fix, nothing to copy.
hq ADR 0107 / issue 115. HAL's own postgres and lavinmq both used a
directory bind (./db-data, ./data) for exactly this data -- the mesh's
adoption of them, three weeks ago, switched to a named Docker volume
instead, and searxng/distribution followed the same pattern since.
A named volume survives ordinary container recreation, same as a
directory bind -- that was never the problem. The problem is everything
else: docker rm -fv, docker volume rm, and docker system prune --volumes
all target it (one flag away from the docker rm -f this migration already
uses routinely); it is invisible to every tool this migration has used all
night (ls, find, grep across /var/lib, /services); and nothing outside
Docker's own volume machinery can back it up or notice it growing.
mesh-store carries the sharpest version: every database migrated tonight,
including keycloak's, live inside it.
Data already copied and verified on novox before this merges:
- mesh-registry-data -> /var/lib/mesh-registry (11G, diff -rq clean)
- mesh-broker-data -> /var/lib/mesh-broker (37M, cp -a)
- mesh-broker-tls -> /var/lib/mesh-broker-tls (12K, cp -a)
- mesh-store-data -> /var/lib/mesh-store (1.7G) -- mesh-store stopped
cleanly first, so the final copy is crash-consistent, not a live-file
copy of a running postgres; diff -rq clean after.
- searxng/valkey: not yet assigned anywhere, manifest-only fix, nothing
to copy.
Old named volumes left in place, not deleted, as the rollback path.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
hq ADR 0107 / issue 115. HAL's own
postgresandlavinmqboth used a directory bind (./db-data,./data) for exactly this data — the mesh's adoption of them three weeks ago switched to a named Docker volume instead, anddistribution/searxngfollowed the same pattern since.Already live on
novox, not just proposed — this was urgent enough (operator instruction, mid-session) to build and roll out before opening this PR. Full account of what happened:mesh-registry-data→/var/lib/mesh-registry(11G,diff -rqclean),mesh-broker-data/mesh-broker-tls→/var/lib/mesh-broker(-tls)(37M + 12K),mesh-store-data→/var/lib/mesh-store(1.7G) —mesh-storestopped cleanly first so the final copy is crash-consistent, not a live-file copy of a running postgres;diff -rqclean after.sudo, somesh-storecrash-looped onmkdir: cannot create directory '/var/lib/postgresql/data': Permission denied— a host bind mount doesn't get auto-chowned by Docker the way a named volume does. Fixed by matching the original volume's ownership (70:70, mode1777) onto the new directory.mesh-brokerandmesh-registrydidn't hit this — their copied files' ownership was already compatible. Totalmesh-storedowntime: a few minutes, during whichmesh-controlleritself couldn't reach its own inventory (expected — it's also onmesh-store) and briefly restart-looped; recovered on its own oncemesh-storecame back.mesh_novox_gitea,mesh_novox_umami,mesh_novox_catalog,mesh_novox_keycloak),keycloak's realm/user counts unchanged (2 realms, 13 users, matching the count from the restore earlier tonight),mesh-brokerhealthy with active connections,mesh-registryserving,gitea's sidecar self-healed.searxng/valkeyisn't assigned anywhere yet — manifest-only fix, nothing to copy.