Data outlives the mesh that declared it, and the conversion starts where it lives
0030, found by asking what the conversion actually needs rather than by reviewing anything. The host deleted a directory and everything under it when it stopped being declared — which happens when a module is unassigned, or when a manifest is edited to move a data folder, which is the exact operation this plan needs. A database's files, a mail spool. The report said "removed". A directory still holding something is now kept and said so. No flag and nothing to remember: emptiness is the test, and it works because the removal order was already right — the mesh's own contents are gone by the time the directory is reached, so what remains is by definition something nobody declared. The plan now says data outranks its own ordering: copy, read back through the service that owns it, and only then point anything at the new location. Never move and then check. And it records where this starts — the node holding all the production data — with what that costs stated rather than argued with. Everything proven so far was proven on machines that could be destroyed and raised again. A scenario proves the mechanism, not the state on that machine.
This commit is contained in:
@@ -115,13 +115,47 @@ endpoint, so *a consumer cannot reach another consumer's bucket* is a policy som
|
||||
a policy granting everything would pass every other test. **What is asserted is what the policy
|
||||
does not say.**
|
||||
|
||||
## Data is the constraint, and it outranks the order below
|
||||
|
||||
*2026-08-31.* The modules being converted run live services — identity, mail — and **the data must
|
||||
survive every step**. A data folder may move; it may never be lost.
|
||||
|
||||
**One thing was found by asking this and is fixed**
|
||||
([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)): the host deleted
|
||||
a directory and everything under it when the directory stopped being declared, which happens when
|
||||
a module is unassigned or a manifest is edited to move a data folder — the exact operation this
|
||||
plan needs. A directory holding anything the mesh did not put there is now kept and reported.
|
||||
|
||||
**That is not a backup and must not be read as one.** It stops the mesh destroying data. It does
|
||||
nothing about a disk, a mistaken command, or a service corrupting its own store.
|
||||
|
||||
**So the rule for every step below:** the data is copied, the copy is verified by reading it back
|
||||
through the service that owns it, and only then does anything point at the new location. Never
|
||||
moved and then checked. **A backup nobody has restored is a belief, not a copy.**
|
||||
|
||||
## Where it starts, and what that costs
|
||||
|
||||
**On the node holding all the production data**, because that is where the services being
|
||||
converted actually are.
|
||||
|
||||
Recorded plainly rather than argued with: this is the highest-risk order available. Everything
|
||||
proven so far was proven on machines that could be destroyed and raised again, and the first real
|
||||
exercise of the conversion will be on the one machine where a mistake is not recoverable. Nothing
|
||||
about the lab work transfers automatically — a scenario proves the mechanism, not the state on
|
||||
that machine.
|
||||
|
||||
**What makes it survivable is preparation rather than caution**: a restored backup before the
|
||||
first step, one service at a time, and the previous arrangement left standing until the new one
|
||||
has been read back. None of that is slower than the alternative, because the alternative includes
|
||||
losing something.
|
||||
|
||||
## Phase 2 — the first real module
|
||||
|
||||
| # | task | done when |
|
||||
|---|---|---|
|
||||
| 2.1 | Port an **object store** module | it runs on the new mesh, serves a bucket to another module, and its credential rotates |
|
||||
| 2.2 | Run it beside the existing one | both exist; nothing depends on the new one yet |
|
||||
| 2.3 | Move one dependent onto it | something real reads and writes through the new mesh's copy |
|
||||
| 2.2 | Copy the data, and read it back through the service that owns it | the new location answers with what the old one holds |
|
||||
| 2.3 | Point one dependent at it, old arrangement left standing | something real reads and writes through the new mesh's copy |
|
||||
|
||||
**Checkpoint, and it is a human one:** it runs for a week before anything else moves. The point of
|
||||
going first is to find what Phase 1 missed, and a week is roughly how long that takes to show.
|
||||
|
||||
Reference in New Issue
Block a user