diff --git a/02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md b/02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md new file mode 100644 index 0000000..a792ee8 --- /dev/null +++ b/02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md @@ -0,0 +1,96 @@ +--- +topic: the tiers +status: accepted +date: 2026-08-31 +deciders: jochen +reconstructed: false +extends: 02-DECISIONS/0005-the-node-host.md +--- + +# 30. Data outlives the mesh that declared it + +## Context + +**The conversion runs on live services holding real data**, and starts on the node that holds all +of it. Identity, mail, everything. The requirement stated plainly: a data directory may be +*moved*, and may never be *lost*. + +**The host deleted them.** A directory that stopped being declared was an orphan, and an orphan +directory was removed with `os.RemoveAll` — everything under it — while the report said +`removed`. A module unassigned took its database's files with it, and nothing anywhere said what +had been in there. + +Reproduced before it was fixed: assign a module, let a service write into its directory, unassign +the module, and the file is gone. + +**A directory stops being declared for ordinary reasons**, which is what makes this sharp rather +than theoretical. A module unassigned from a node. A manifest edited to move a data folder — the +exact operation the conversion needs. A resource renamed. A typo. **Every one of those is a normal +day's work, and every one of them was destructive.** + +**The removal order was already right, and that is what makes a fix possible.** Everything the +mesh puts inside a directory is itself a declared resource, and orphans are removed in reverse +declaration order — so by the time a directory is reached, what the mesh wrote there is already +gone. Anything still present was put there by something else. + +## Considered Options + +1. **A `keep` flag on the directory.** A module declares which of its directories hold data, and + the host leaves those. **Rejected.** It is safe only when somebody remembered, and the failure + of forgetting is total and silent. A rule that protects data only when it was asked to is not + a rule about data, it is a rule about attentiveness — and this is the one place in the system + where being wrong does not recover. + +2. **Never remove a directory.** Simple and unarguably safe. **Rejected**, narrowly: every module + ever assigned would leave its directories behind for ever, and a machine that accumulates + things nobody can account for is one where nobody can tell what is still in use. The clean-up + that is genuinely the mesh's is worth keeping. + +3. **Remove a directory only when it is empty.** **Adopted.** + +## Decision + +**A directory that still holds anything is kept, and the mesh says so.** An empty one is removed. + +**This is the host's existing line applied to the one shape where getting it wrong is +unrecoverable** — *it removes what it made and leaves what it merely configured* +([ADR 0005](0005-the-node-host.md)). An empty directory is what the host made. A full one is not. + +**No flag, no declaration, nothing to remember.** Emptiness is the test, and it is derived from +the removal order rather than asserted: the mesh's own contents are gone by then, so what remains +is by definition something nobody declared. + +**It is reported, not silent.** The outcome is `kept`, naming how many items are inside and saying +they are for a person to deal with. A directory quietly left behind is how a machine accumulates +things nobody can account for — which is the objection to option 2, and it is answered by saying +so rather than by deleting. + +**Files are unchanged.** A declared file is the mesh's own — it wrote it, it owns it, and losing a +configuration file is not the failure this is about. The distinction is deliberate: **directories +hold what other things produced; files are what the mesh itself put there.** + +## Consequences + +**Moving a data directory is now safe by default.** The manifest changes, the old path stops being +declared, and the data stays where it is until somebody has looked at it. That was the operation +most likely to destroy something during the conversion, and it is now the operation that does the +least. + +**Unassigning a module leaves its data.** Correct, and it means unassignment is no longer a way to +clean up — removing data is a person's act, done knowingly. Given what unassignment did before, +that is the trade being made and it is the right way round. + +**A machine can accumulate directories nobody removed.** Accepted, and mitigated by saying so +every time rather than by a periodic sweep. A sweep would be the deletion this record exists to +prevent, on a timer, with nobody watching. + +**It is not a backup, and must not be mistaken for one.** This stops the mesh destroying data. It +does nothing about a disk, a mistaken `rm`, or a service corrupting its own store. The conversion +still needs backups taken and **restored** before anything is moved — a backup nobody has restored +is a belief, not a copy. + +## References + +- [ADR 0005](0005-the-node-host.md) — the host removes what it made +- [`03-DESIGN/01-to-be/00-work-breakdown.md`](../03-DESIGN/01-to-be/00-work-breakdown.md) — the + conversion this was found by planning diff --git a/02-DECISIONS/README.md b/02-DECISIONS/README.md index 6302c44..6a45ec7 100644 --- a/02-DECISIONS/README.md +++ b/02-DECISIONS/README.md @@ -94,6 +94,7 @@ python3 00-META/checks/index.py fail if stale - **0008** — [A context owns its store, exclusively](0008-a-context-owns-its-store.md) - **0028** — [The substrate supplies the control plane and nothing else](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md) - **0029** — [A network is a shape, because an action cannot be undone](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md) +- **0030** — [Data outlives the mesh that declared it](0030-data-outlives-the-mesh-that-declared-it.md) ### What runs on them, and how it gets there diff --git a/03-DESIGN/01-to-be/00-work-breakdown.md b/03-DESIGN/01-to-be/00-work-breakdown.md index 7cb1622..e458d96 100644 --- a/03-DESIGN/01-to-be/00-work-breakdown.md +++ b/03-DESIGN/01-to-be/00-work-breakdown.md @@ -115,13 +115,47 @@ endpoint, so *a consumer cannot reach another consumer's bucket* is a policy som a policy granting everything would pass every other test. **What is asserted is what the policy does not say.** +## Data is the constraint, and it outranks the order below + +*2026-08-31.* The modules being converted run live services — identity, mail — and **the data must +survive every step**. A data folder may move; it may never be lost. + +**One thing was found by asking this and is fixed** +([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)): the host deleted +a directory and everything under it when the directory stopped being declared, which happens when +a module is unassigned or a manifest is edited to move a data folder — the exact operation this +plan needs. A directory holding anything the mesh did not put there is now kept and reported. + +**That is not a backup and must not be read as one.** It stops the mesh destroying data. It does +nothing about a disk, a mistaken command, or a service corrupting its own store. + +**So the rule for every step below:** the data is copied, the copy is verified by reading it back +through the service that owns it, and only then does anything point at the new location. Never +moved and then checked. **A backup nobody has restored is a belief, not a copy.** + +## Where it starts, and what that costs + +**On the node holding all the production data**, because that is where the services being +converted actually are. + +Recorded plainly rather than argued with: this is the highest-risk order available. Everything +proven so far was proven on machines that could be destroyed and raised again, and the first real +exercise of the conversion will be on the one machine where a mistake is not recoverable. Nothing +about the lab work transfers automatically — a scenario proves the mechanism, not the state on +that machine. + +**What makes it survivable is preparation rather than caution**: a restored backup before the +first step, one service at a time, and the previous arrangement left standing until the new one +has been read back. None of that is slower than the alternative, because the alternative includes +losing something. + ## Phase 2 — the first real module | # | task | done when | |---|---|---| | 2.1 | Port an **object store** module | it runs on the new mesh, serves a bucket to another module, and its credential rotates | -| 2.2 | Run it beside the existing one | both exist; nothing depends on the new one yet | -| 2.3 | Move one dependent onto it | something real reads and writes through the new mesh's copy | +| 2.2 | Copy the data, and read it back through the service that owns it | the new location answers with what the old one holds | +| 2.3 | Point one dependent at it, old arrangement left standing | something real reads and writes through the new mesh's copy | **Checkpoint, and it is a human one:** it runs for a week before anything else moves. The point of going first is to find what Phase 1 missed, and a week is roughly how long that takes to show.