Data outlives the mesh that declared it, and the conversion starts where it lives

0030, found by asking what the conversion actually needs rather than by
reviewing anything. The host deleted a directory and everything under it
when it stopped being declared — which happens when a module is
unassigned, or when a manifest is edited to move a data folder, which is
the exact operation this plan needs. A database's files, a mail spool.
The report said "removed".

A directory still holding something is now kept and said so. No flag and
nothing to remember: emptiness is the test, and it works because the
removal order was already right — the mesh's own contents are gone by
the time the directory is reached, so what remains is by definition
something nobody declared.

The plan now says data outranks its own ordering: copy, read back
through the service that owns it, and only then point anything at the
new location. Never move and then check.

And it records where this starts — the node holding all the production
data — with what that costs stated rather than argued with. Everything
proven so far was proven on machines that could be destroyed and raised
again. A scenario proves the mechanism, not the state on that machine.
This commit is contained in:
2026-08-31 19:54:41 +02:00
parent 9ad0ec35e0
commit d1ab2dc0b4
3 changed files with 133 additions and 2 deletions
+36 -2
View File
@@ -115,13 +115,47 @@ endpoint, so *a consumer cannot reach another consumer's bucket* is a policy som
a policy granting everything would pass every other test. **What is asserted is what the policy
does not say.**
## Data is the constraint, and it outranks the order below
*2026-08-31.* The modules being converted run live services — identity, mail — and **the data must
survive every step**. A data folder may move; it may never be lost.
**One thing was found by asking this and is fixed**
([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)): the host deleted
a directory and everything under it when the directory stopped being declared, which happens when
a module is unassigned or a manifest is edited to move a data folder — the exact operation this
plan needs. A directory holding anything the mesh did not put there is now kept and reported.
**That is not a backup and must not be read as one.** It stops the mesh destroying data. It does
nothing about a disk, a mistaken command, or a service corrupting its own store.
**So the rule for every step below:** the data is copied, the copy is verified by reading it back
through the service that owns it, and only then does anything point at the new location. Never
moved and then checked. **A backup nobody has restored is a belief, not a copy.**
## Where it starts, and what that costs
**On the node holding all the production data**, because that is where the services being
converted actually are.
Recorded plainly rather than argued with: this is the highest-risk order available. Everything
proven so far was proven on machines that could be destroyed and raised again, and the first real
exercise of the conversion will be on the one machine where a mistake is not recoverable. Nothing
about the lab work transfers automatically — a scenario proves the mechanism, not the state on
that machine.
**What makes it survivable is preparation rather than caution**: a restored backup before the
first step, one service at a time, and the previous arrangement left standing until the new one
has been read back. None of that is slower than the alternative, because the alternative includes
losing something.
## Phase 2 — the first real module
| # | task | done when |
|---|---|---|
| 2.1 | Port an **object store** module | it runs on the new mesh, serves a bucket to another module, and its credential rotates |
| 2.2 | Run it beside the existing one | both exist; nothing depends on the new one yet |
| 2.3 | Move one dependent onto it | something real reads and writes through the new mesh's copy |
| 2.2 | Copy the data, and read it back through the service that owns it | the new location answers with what the old one holds |
| 2.3 | Point one dependent at it, old arrangement left standing | something real reads and writes through the new mesh's copy |
**Checkpoint, and it is a human one:** it runs for a week before anything else moves. The point of
going first is to find what Phase 1 missed, and a week is roughly how long that takes to show.