ADR 0107: persistent data is a directory bind, never a named volume

Records the rule the operator gave directly, mid-session, after checking
that HAL's own postgres and lavinmq both used a directory bind and the
mesh's adoption of them three weeks ago switched to a named volume without
a reason recorded anywhere.

Already built and rolled out on novox (mesh-catalog PR #54) before this
record -- urgent enough to fix first and write down after. Includes the
incident: the new host directories needed the container's own UID, which a
named volume gets for free and a directory bind does not; mesh-store
crash-looped on Permission denied until ownership was matched to what the
original volume already had.

Closes issue 115. Checks pass.
This commit is contained in:
2026-09-24 16:21:51 +02:00
parent a458f751c6
commit a93743708c
3 changed files with 147 additions and 0 deletions
@@ -0,0 +1,71 @@
---
topic: building it
status: accepted
date: 2026-09-24
deciders: jochen
reconstructed: false
---
# 107. Persistent data is a directory bind, never a named volume
## Context
Measured 2026-09-24, mid-migration: four modules mount a named Docker volume for real state —
`mesh-store` (every database the mesh holds), `mesh-broker` (its data and TLS material),
`mesh-registry` (every image), and `searxng`'s cache. Every other module in the catalogue —
more than forty — mounts a host directory, `/var/lib/<module>/...`.
The predecessor did not make this choice. HAL's own `postgres` bound `./db-data`, and its
`lavinmq` bound `./data` — directories, both. The mesh's adoption of them
([`a63ef3d`](https://git.novox.be/novox/mesh-catalog/commit/a63ef3d), "the postgres module
adopts mesh-store instead of raising its own", 2026-09-16) introduced the named volume; three
other modules followed the same shape since. Full account of what was found and fixed:
[issue 115](../04-ISSUES/115-a-named-docker-volume-is-invisible-and-one-flag-from-gone/00-report.md).
## The two guarantees are not the same guarantee
A named volume and a host directory both survive **ordinary** container recreation — a rebuild,
a `take`, the `docker rm -f` and push this migration already uses routinely. Neither loses data
to that. That was never the question.
What they do not both survive:
- **`docker rm -fv`, `docker volume rm`, `docker system prune --volumes`** all target a named
volume specifically. The first is one character from the command this migration's own rules
already call for after every address change. A host directory has no equivalent command that
destroys it by accident — removing it is always a deliberate `rm -rf` on a path someone typed.
- **Visibility.** Every tool this migration has used all night to find and verify data —
`ls`, `find`, `grep`, a backup job — reaches a host directory for free. A named volume requires
knowing to ask Docker (`docker volume inspect`) before its bytes, at
`/var/lib/docker/volumes/<name>/_data`, are reachable at all.
## Decision
**A container mount holding data that must survive is a host directory bind. A named volume is
permitted only for data that is disposable if lost** — a cache, a scratch space, something the
module rebuilds on next start without consequence. `searxng`'s `valkey` cache is close to this
line and was converted anyway, for consistency and because it costs nothing to.
Ownership is the one thing a host directory does not get for free that a named volume does:
Docker initialises a fresh named volume's ownership to what the container's first process needs;
a host directory is whatever created it. A directory made for this purpose must be given the
image's expected UID before the container using it starts — read from the running instance being
replaced when one exists, rather than guessed.
## Consequences
- The four modules were converted: `distribution`, `lavinmq`, `postgres`, `searxng`. Data copied
and verified byte-for-byte before each manifest changed; `mesh-store` stopped cleanly first, so
its copy is crash-consistent rather than a live read of a running postgres.
- **The ownership gap above was not theoretical — it is what happened.** The new directories
were created by the operator's tooling running as root; `mesh-store` crash-looped on
`mkdir: ... Permission denied` until its directory's ownership was set to match what the
original volume already had. Worth a check at `module add` time — nothing catches this today
beyond the container failing to start.
- Old named volumes were not deleted. They remain the rollback path until confidence in the new
mounts is established over time, not one clean start.
## References
- [Issue 115](../04-ISSUES/115-a-named-docker-volume-is-invisible-and-one-flag-from-gone/00-report.md)
- `mesh-catalog` PR #54
+1
View File
@@ -172,6 +172,7 @@ python3 00-META/checks/index.py fail if stale
- **0086** — [A secret reaches a process as a file, and an exception is declared](0086-a-secret-reaches-a-process-as-a-file.md)
- **0096** — [An upstream image is copied between registries, never through a machine's image store](0096-an-upstream-image-is-copied-between-registries.md)
- **0097** — [A vendor image is a declared build input, and a recipe fetches nothing undeclared](0097-a-vendor-image-is-a-declared-build-input.md)
- **0107** — [Persistent data is a directory bind, never a named volume](0107-persistent-data-is-a-directory-bind-never-a-named-volume.md)
### How it is checked