Records the rule the operator gave directly, mid-session, after checking that HAL's own postgres and lavinmq both used a directory bind and the mesh's adoption of them three weeks ago switched to a named volume without a reason recorded anywhere. Already built and rolled out on novox (mesh-catalog PR #54) before this record -- urgent enough to fix first and write down after. Includes the incident: the new host directories needed the container's own UID, which a named volume gets for free and a directory bind does not; mesh-store crash-looped on Permission denied until ownership was matched to what the original volume already had. Closes issue 115. Checks pass.
76 lines
4.2 KiB
Markdown
76 lines
4.2 KiB
Markdown
---
|
|
status: resolved
|
|
opened: 2026-09-24
|
|
located-in: [mesh-catalog]
|
|
fixed-by: mesh-catalog PR #54 — distribution, lavinmq, postgres and searxng converted to host directory binds; data copied and verified (mesh-store stopped cleanly first for a crash-consistent copy), old named volumes kept as the rollback path
|
|
amended-design:
|
|
---
|
|
|
|
# 115 — A named Docker volume is invisible to the operator, and one flag from gone
|
|
|
|
## What was observed
|
|
|
|
On the control-node, 2026-09-24, mid-migration, checking every module in the catalogue for how it
|
|
mounts its data. Five container mounts across four modules use a **named Docker volume** rather
|
|
than a host directory:
|
|
|
|
```
|
|
distribution mesh-registry → mesh-registry-data:/var/lib/registry
|
|
lavinmq mesh-broker → mesh-broker-data:/var/lib/lavinmq (+ mesh-broker-tls)
|
|
postgres mesh-store → mesh-store-data:/var/lib/postgresql/data
|
|
searxng valkey → searxng-valkey-data:/data
|
|
```
|
|
|
|
Every other module in the catalogue — more than forty of them — mounts a host directory,
|
|
`/var/lib/<module>/...`, matching what `DATA-CUTOVER.md` and every rehearsed recipe tonight
|
|
assumes. These four are the exception, not a second convention.
|
|
|
|
`mesh-store` is the one that matters most: it holds every database migrated tonight, including a
|
|
live `keycloak` restore verified minutes before this was written.
|
|
|
|
## Why it matters beyond this instance
|
|
|
|
[`03-DESIGN/01-to-be/22-the-work-ahead.md`](../../03-DESIGN/01-to-be/22-the-work-ahead.md) shows
|
|
this was a deliberate choice, not an oversight — *"Data survives on the named volumes"* — but that
|
|
sentence answers a narrower question than the one this issue raises. It says a named volume
|
|
survives **ordinary container recreation** (a rebuild, a `take`, a routine `docker rm -f` and
|
|
push), which is true and which every module already gets from either a named volume or a host
|
|
directory equally.
|
|
|
|
What it does not address: a named volume is **one flag away from deleted**, in a way a host
|
|
directory structurally cannot be.
|
|
|
|
- `docker rm -f` alone does not remove a named volume — it persists, unreferenced, until something
|
|
targets it by name.
|
|
- `docker rm -fv`, `docker volume rm`, and `docker system prune --volumes` all do target it, and
|
|
the difference from the command this migration already uses routinely (`docker rm -f` — see
|
|
`HANDOFF.md`'s own "a restart does not re-read anything" rule) is one character.
|
|
- A named volume is invisible to an operator working the way this migration has worked all
|
|
night: `ls`, `find`, `grep` across `/services/*` and `/var/lib/*`. Finding it requires knowing to
|
|
ask Docker (`docker volume inspect`), and its actual bytes sit under
|
|
`/var/lib/docker/volumes/<name>/_data`, a path nothing points at.
|
|
- Nothing external can back it up, snapshot it, or notice it growing without going through
|
|
Docker's own volume machinery — a host directory is a directory; a filesystem-level backup job
|
|
already reaches it for free.
|
|
|
|
`mesh-store` carries the sharpest version of this: every module's database, the mesh's own
|
|
inventory, identity and licence stores — the single foundation piece the rest of the mesh depends
|
|
on — sits somewhere the operator's ordinary tools do not look.
|
|
|
|
## Decided
|
|
|
|
**Persistent data is a host directory bind, never a named volume. A named volume may hold only
|
|
data that is disposable if lost.** All four hold real state and all four are in scope — including
|
|
`mesh-registry` and `searxng`'s cache, not only `mesh-store`. Decided 2026-09-24; the record is
|
|
[ADR 0107](../../02-DECISIONS/0107-persistent-data-is-a-directory-bind-never-a-named-volume.md).
|
|
|
|
## Open questions
|
|
|
|
- The safe migration path for each, in order of stakes — `mesh-store` live and holding every
|
|
database migrated tonight, `mesh-broker`, `mesh-registry`, then `searxng`'s cache, which is
|
|
genuinely disposable and may not need migrating at all if it is rebuilt rather than moved.
|
|
- Should the catalogue refuse a module declaring a named volume for anything but disposable data
|
|
at `module add`, the way [issue 091](../091-a-module-definition-carries-a-machine-port/00-report.md)
|
|
asks the same of a hardcoded machine port — so a convention violation is caught at registration
|
|
rather than found by reading the whole catalogue?
|