Files
hq/04-ISSUES/115-a-named-docker-volume-is-invisible-and-one-flag-from-gone/00-report.md
T
jschoubben a93743708c ADR 0107: persistent data is a directory bind, never a named volume
Records the rule the operator gave directly, mid-session, after checking
that HAL's own postgres and lavinmq both used a directory bind and the
mesh's adoption of them three weeks ago switched to a named volume without
a reason recorded anywhere.

Already built and rolled out on novox (mesh-catalog PR #54) before this
record -- urgent enough to fix first and write down after. Includes the
incident: the new host directories needed the container's own UID, which a
named volume gets for free and a directory bind does not; mesh-store
crash-looped on Permission denied until ownership was matched to what the
original volume already had.

Closes issue 115. Checks pass.
2026-09-24 16:21:51 +02:00

4.2 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
resolved 2026-09-24
mesh-catalog
mesh-catalog PR

115 — A named Docker volume is invisible to the operator, and one flag from gone

What was observed

On the control-node, 2026-09-24, mid-migration, checking every module in the catalogue for how it mounts its data. Five container mounts across four modules use a named Docker volume rather than a host directory:

distribution   mesh-registry  →  mesh-registry-data:/var/lib/registry
lavinmq        mesh-broker    →  mesh-broker-data:/var/lib/lavinmq  (+ mesh-broker-tls)
postgres       mesh-store     →  mesh-store-data:/var/lib/postgresql/data
searxng        valkey         →  searxng-valkey-data:/data

Every other module in the catalogue — more than forty of them — mounts a host directory, /var/lib/<module>/..., matching what DATA-CUTOVER.md and every rehearsed recipe tonight assumes. These four are the exception, not a second convention.

mesh-store is the one that matters most: it holds every database migrated tonight, including a live keycloak restore verified minutes before this was written.

Why it matters beyond this instance

03-DESIGN/01-to-be/22-the-work-ahead.md shows this was a deliberate choice, not an oversight — "Data survives on the named volumes" — but that sentence answers a narrower question than the one this issue raises. It says a named volume survives ordinary container recreation (a rebuild, a take, a routine docker rm -f and push), which is true and which every module already gets from either a named volume or a host directory equally.

What it does not address: a named volume is one flag away from deleted, in a way a host directory structurally cannot be.

  • docker rm -f alone does not remove a named volume — it persists, unreferenced, until something targets it by name.
  • docker rm -fv, docker volume rm, and docker system prune --volumes all do target it, and the difference from the command this migration already uses routinely (docker rm -f — see HANDOFF.md's own "a restart does not re-read anything" rule) is one character.
  • A named volume is invisible to an operator working the way this migration has worked all night: ls, find, grep across /services/* and /var/lib/*. Finding it requires knowing to ask Docker (docker volume inspect), and its actual bytes sit under /var/lib/docker/volumes/<name>/_data, a path nothing points at.
  • Nothing external can back it up, snapshot it, or notice it growing without going through Docker's own volume machinery — a host directory is a directory; a filesystem-level backup job already reaches it for free.

mesh-store carries the sharpest version of this: every module's database, the mesh's own inventory, identity and licence stores — the single foundation piece the rest of the mesh depends on — sits somewhere the operator's ordinary tools do not look.

Decided

Persistent data is a host directory bind, never a named volume. A named volume may hold only data that is disposable if lost. All four hold real state and all four are in scope — including mesh-registry and searxng's cache, not only mesh-store. Decided 2026-09-24; the record is ADR 0107.

Open questions

  • The safe migration path for each, in order of stakes — mesh-store live and holding every database migrated tonight, mesh-broker, mesh-registry, then searxng's cache, which is genuinely disposable and may not need migrating at all if it is rebuilt rather than moved.
  • Should the catalogue refuse a module declaring a named volume for anything but disposable data at module add, the way issue 091 asks the same of a hardcoded machine port — so a convention violation is caught at registration rather than found by reading the whole catalogue?