Files
hq/04-ISSUES/115-a-named-docker-volume-is-invisible-and-one-flag-from-gone/00-report.md
T
jschoubben a93743708c ADR 0107: persistent data is a directory bind, never a named volume
Records the rule the operator gave directly, mid-session, after checking
that HAL's own postgres and lavinmq both used a directory bind and the
mesh's adoption of them three weeks ago switched to a named volume without
a reason recorded anywhere.

Already built and rolled out on novox (mesh-catalog PR #54) before this
record -- urgent enough to fix first and write down after. Includes the
incident: the new host directories needed the container's own UID, which a
named volume gets for free and a directory bind does not; mesh-store
crash-looped on Permission denied until ownership was matched to what the
original volume already had.

Closes issue 115. Checks pass.
2026-09-24 16:21:51 +02:00

76 lines
4.2 KiB
Markdown

---
status: resolved
opened: 2026-09-24
located-in: [mesh-catalog]
fixed-by: mesh-catalog PR #54 — distribution, lavinmq, postgres and searxng converted to host directory binds; data copied and verified (mesh-store stopped cleanly first for a crash-consistent copy), old named volumes kept as the rollback path
amended-design:
---
# 115 — A named Docker volume is invisible to the operator, and one flag from gone
## What was observed
On the control-node, 2026-09-24, mid-migration, checking every module in the catalogue for how it
mounts its data. Five container mounts across four modules use a **named Docker volume** rather
than a host directory:
```
distribution mesh-registry → mesh-registry-data:/var/lib/registry
lavinmq mesh-broker → mesh-broker-data:/var/lib/lavinmq (+ mesh-broker-tls)
postgres mesh-store → mesh-store-data:/var/lib/postgresql/data
searxng valkey → searxng-valkey-data:/data
```
Every other module in the catalogue — more than forty of them — mounts a host directory,
`/var/lib/<module>/...`, matching what `DATA-CUTOVER.md` and every rehearsed recipe tonight
assumes. These four are the exception, not a second convention.
`mesh-store` is the one that matters most: it holds every database migrated tonight, including a
live `keycloak` restore verified minutes before this was written.
## Why it matters beyond this instance
[`03-DESIGN/01-to-be/22-the-work-ahead.md`](../../03-DESIGN/01-to-be/22-the-work-ahead.md) shows
this was a deliberate choice, not an oversight — *"Data survives on the named volumes"* — but that
sentence answers a narrower question than the one this issue raises. It says a named volume
survives **ordinary container recreation** (a rebuild, a `take`, a routine `docker rm -f` and
push), which is true and which every module already gets from either a named volume or a host
directory equally.
What it does not address: a named volume is **one flag away from deleted**, in a way a host
directory structurally cannot be.
- `docker rm -f` alone does not remove a named volume — it persists, unreferenced, until something
targets it by name.
- `docker rm -fv`, `docker volume rm`, and `docker system prune --volumes` all do target it, and
the difference from the command this migration already uses routinely (`docker rm -f` — see
`HANDOFF.md`'s own "a restart does not re-read anything" rule) is one character.
- A named volume is invisible to an operator working the way this migration has worked all
night: `ls`, `find`, `grep` across `/services/*` and `/var/lib/*`. Finding it requires knowing to
ask Docker (`docker volume inspect`), and its actual bytes sit under
`/var/lib/docker/volumes/<name>/_data`, a path nothing points at.
- Nothing external can back it up, snapshot it, or notice it growing without going through
Docker's own volume machinery — a host directory is a directory; a filesystem-level backup job
already reaches it for free.
`mesh-store` carries the sharpest version of this: every module's database, the mesh's own
inventory, identity and licence stores — the single foundation piece the rest of the mesh depends
on — sits somewhere the operator's ordinary tools do not look.
## Decided
**Persistent data is a host directory bind, never a named volume. A named volume may hold only
data that is disposable if lost.** All four hold real state and all four are in scope — including
`mesh-registry` and `searxng`'s cache, not only `mesh-store`. Decided 2026-09-24; the record is
[ADR 0107](../../02-DECISIONS/0107-persistent-data-is-a-directory-bind-never-a-named-volume.md).
## Open questions
- The safe migration path for each, in order of stakes — `mesh-store` live and holding every
database migrated tonight, `mesh-broker`, `mesh-registry`, then `searxng`'s cache, which is
genuinely disposable and may not need migrating at all if it is rebuilt rather than moved.
- Should the catalogue refuse a module declaring a named volume for anything but disposable data
at `module add`, the way [issue 091](../091-a-module-definition-carries-a-machine-port/00-report.md)
asks the same of a hardcoded machine port — so a convention violation is caught at registration
rather than found by reading the whole catalogue?