diff --git a/02-DECISIONS/0107-persistent-data-is-a-directory-bind-never-a-named-volume.md b/02-DECISIONS/0107-persistent-data-is-a-directory-bind-never-a-named-volume.md new file mode 100644 index 0000000..2054454 --- /dev/null +++ b/02-DECISIONS/0107-persistent-data-is-a-directory-bind-never-a-named-volume.md @@ -0,0 +1,71 @@ +--- +topic: building it +status: accepted +date: 2026-09-24 +deciders: jochen +reconstructed: false +--- + +# 107. Persistent data is a directory bind, never a named volume + +## Context + +Measured 2026-09-24, mid-migration: four modules mount a named Docker volume for real state — +`mesh-store` (every database the mesh holds), `mesh-broker` (its data and TLS material), +`mesh-registry` (every image), and `searxng`'s cache. Every other module in the catalogue — +more than forty — mounts a host directory, `/var/lib//...`. + +The predecessor did not make this choice. HAL's own `postgres` bound `./db-data`, and its +`lavinmq` bound `./data` — directories, both. The mesh's adoption of them +([`a63ef3d`](https://git.novox.be/novox/mesh-catalog/commit/a63ef3d), "the postgres module +adopts mesh-store instead of raising its own", 2026-09-16) introduced the named volume; three +other modules followed the same shape since. Full account of what was found and fixed: +[issue 115](../04-ISSUES/115-a-named-docker-volume-is-invisible-and-one-flag-from-gone/00-report.md). + +## The two guarantees are not the same guarantee + +A named volume and a host directory both survive **ordinary** container recreation — a rebuild, +a `take`, the `docker rm -f` and push this migration already uses routinely. Neither loses data +to that. That was never the question. + +What they do not both survive: + +- **`docker rm -fv`, `docker volume rm`, `docker system prune --volumes`** all target a named + volume specifically. The first is one character from the command this migration's own rules + already call for after every address change. A host directory has no equivalent command that + destroys it by accident — removing it is always a deliberate `rm -rf` on a path someone typed. +- **Visibility.** Every tool this migration has used all night to find and verify data — + `ls`, `find`, `grep`, a backup job — reaches a host directory for free. A named volume requires + knowing to ask Docker (`docker volume inspect`) before its bytes, at + `/var/lib/docker/volumes//_data`, are reachable at all. + +## Decision + +**A container mount holding data that must survive is a host directory bind. A named volume is +permitted only for data that is disposable if lost** — a cache, a scratch space, something the +module rebuilds on next start without consequence. `searxng`'s `valkey` cache is close to this +line and was converted anyway, for consistency and because it costs nothing to. + +Ownership is the one thing a host directory does not get for free that a named volume does: +Docker initialises a fresh named volume's ownership to what the container's first process needs; +a host directory is whatever created it. A directory made for this purpose must be given the +image's expected UID before the container using it starts — read from the running instance being +replaced when one exists, rather than guessed. + +## Consequences + +- The four modules were converted: `distribution`, `lavinmq`, `postgres`, `searxng`. Data copied + and verified byte-for-byte before each manifest changed; `mesh-store` stopped cleanly first, so + its copy is crash-consistent rather than a live read of a running postgres. +- **The ownership gap above was not theoretical — it is what happened.** The new directories + were created by the operator's tooling running as root; `mesh-store` crash-looped on + `mkdir: ... Permission denied` until its directory's ownership was set to match what the + original volume already had. Worth a check at `module add` time — nothing catches this today + beyond the container failing to start. +- Old named volumes were not deleted. They remain the rollback path until confidence in the new + mounts is established over time, not one clean start. + +## References + +- [Issue 115](../04-ISSUES/115-a-named-docker-volume-is-invisible-and-one-flag-from-gone/00-report.md) +- `mesh-catalog` PR #54 diff --git a/02-DECISIONS/README.md b/02-DECISIONS/README.md index 45cee32..e0ae9bd 100644 --- a/02-DECISIONS/README.md +++ b/02-DECISIONS/README.md @@ -172,6 +172,7 @@ python3 00-META/checks/index.py fail if stale - **0086** — [A secret reaches a process as a file, and an exception is declared](0086-a-secret-reaches-a-process-as-a-file.md) - **0096** — [An upstream image is copied between registries, never through a machine's image store](0096-an-upstream-image-is-copied-between-registries.md) - **0097** — [A vendor image is a declared build input, and a recipe fetches nothing undeclared](0097-a-vendor-image-is-a-declared-build-input.md) +- **0107** — [Persistent data is a directory bind, never a named volume](0107-persistent-data-is-a-directory-bind-never-a-named-volume.md) ### How it is checked diff --git a/04-ISSUES/115-a-named-docker-volume-is-invisible-and-one-flag-from-gone/00-report.md b/04-ISSUES/115-a-named-docker-volume-is-invisible-and-one-flag-from-gone/00-report.md new file mode 100644 index 0000000..0c3a8e7 --- /dev/null +++ b/04-ISSUES/115-a-named-docker-volume-is-invisible-and-one-flag-from-gone/00-report.md @@ -0,0 +1,75 @@ +--- +status: resolved +opened: 2026-09-24 +located-in: [mesh-catalog] +fixed-by: mesh-catalog PR #54 — distribution, lavinmq, postgres and searxng converted to host directory binds; data copied and verified (mesh-store stopped cleanly first for a crash-consistent copy), old named volumes kept as the rollback path +amended-design: +--- + +# 115 — A named Docker volume is invisible to the operator, and one flag from gone + +## What was observed + +On the control-node, 2026-09-24, mid-migration, checking every module in the catalogue for how it +mounts its data. Five container mounts across four modules use a **named Docker volume** rather +than a host directory: + +``` +distribution mesh-registry → mesh-registry-data:/var/lib/registry +lavinmq mesh-broker → mesh-broker-data:/var/lib/lavinmq (+ mesh-broker-tls) +postgres mesh-store → mesh-store-data:/var/lib/postgresql/data +searxng valkey → searxng-valkey-data:/data +``` + +Every other module in the catalogue — more than forty of them — mounts a host directory, +`/var/lib//...`, matching what `DATA-CUTOVER.md` and every rehearsed recipe tonight +assumes. These four are the exception, not a second convention. + +`mesh-store` is the one that matters most: it holds every database migrated tonight, including a +live `keycloak` restore verified minutes before this was written. + +## Why it matters beyond this instance + +[`03-DESIGN/01-to-be/22-the-work-ahead.md`](../../03-DESIGN/01-to-be/22-the-work-ahead.md) shows +this was a deliberate choice, not an oversight — *"Data survives on the named volumes"* — but that +sentence answers a narrower question than the one this issue raises. It says a named volume +survives **ordinary container recreation** (a rebuild, a `take`, a routine `docker rm -f` and +push), which is true and which every module already gets from either a named volume or a host +directory equally. + +What it does not address: a named volume is **one flag away from deleted**, in a way a host +directory structurally cannot be. + +- `docker rm -f` alone does not remove a named volume — it persists, unreferenced, until something + targets it by name. +- `docker rm -fv`, `docker volume rm`, and `docker system prune --volumes` all do target it, and + the difference from the command this migration already uses routinely (`docker rm -f` — see + `HANDOFF.md`'s own "a restart does not re-read anything" rule) is one character. +- A named volume is invisible to an operator working the way this migration has worked all + night: `ls`, `find`, `grep` across `/services/*` and `/var/lib/*`. Finding it requires knowing to + ask Docker (`docker volume inspect`), and its actual bytes sit under + `/var/lib/docker/volumes//_data`, a path nothing points at. +- Nothing external can back it up, snapshot it, or notice it growing without going through + Docker's own volume machinery — a host directory is a directory; a filesystem-level backup job + already reaches it for free. + +`mesh-store` carries the sharpest version of this: every module's database, the mesh's own +inventory, identity and licence stores — the single foundation piece the rest of the mesh depends +on — sits somewhere the operator's ordinary tools do not look. + +## Decided + +**Persistent data is a host directory bind, never a named volume. A named volume may hold only +data that is disposable if lost.** All four hold real state and all four are in scope — including +`mesh-registry` and `searxng`'s cache, not only `mesh-store`. Decided 2026-09-24; the record is +[ADR 0107](../../02-DECISIONS/0107-persistent-data-is-a-directory-bind-never-a-named-volume.md). + +## Open questions + +- The safe migration path for each, in order of stakes — `mesh-store` live and holding every + database migrated tonight, `mesh-broker`, `mesh-registry`, then `searxng`'s cache, which is + genuinely disposable and may not need migrating at all if it is rebuilt rather than moved. +- Should the catalogue refuse a module declaring a named volume for anything but disposable data + at `module add`, the way [issue 091](../091-a-module-definition-carries-a-machine-port/00-report.md) + asks the same of a hardcoded machine port — so a convention violation is caught at registration + rather than found by reading the whole catalogue?