Data outlives the mesh that declared it, and the conversion starts where it lives

0030, found by asking what the conversion actually needs rather than by
reviewing anything. The host deleted a directory and everything under it
when it stopped being declared — which happens when a module is
unassigned, or when a manifest is edited to move a data folder, which is
the exact operation this plan needs. A database's files, a mail spool.
The report said "removed".

A directory still holding something is now kept and said so. No flag and
nothing to remember: emptiness is the test, and it works because the
removal order was already right — the mesh's own contents are gone by
the time the directory is reached, so what remains is by definition
something nobody declared.

The plan now says data outranks its own ordering: copy, read back
through the service that owns it, and only then point anything at the
new location. Never move and then check.

And it records where this starts — the node holding all the production
data — with what that costs stated rather than argued with. Everything
proven so far was proven on machines that could be destroyed and raised
again. A scenario proves the mechanism, not the state on that machine.
This commit is contained in:
2026-08-31 19:54:41 +02:00
parent 9ad0ec35e0
commit d1ab2dc0b4
3 changed files with 133 additions and 2 deletions
@@ -0,0 +1,96 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0005-the-node-host.md
---
# 30. Data outlives the mesh that declared it
## Context
**The conversion runs on live services holding real data**, and starts on the node that holds all
of it. Identity, mail, everything. The requirement stated plainly: a data directory may be
*moved*, and may never be *lost*.
**The host deleted them.** A directory that stopped being declared was an orphan, and an orphan
directory was removed with `os.RemoveAll` — everything under it — while the report said
`removed`. A module unassigned took its database's files with it, and nothing anywhere said what
had been in there.
Reproduced before it was fixed: assign a module, let a service write into its directory, unassign
the module, and the file is gone.
**A directory stops being declared for ordinary reasons**, which is what makes this sharp rather
than theoretical. A module unassigned from a node. A manifest edited to move a data folder — the
exact operation the conversion needs. A resource renamed. A typo. **Every one of those is a normal
day's work, and every one of them was destructive.**
**The removal order was already right, and that is what makes a fix possible.** Everything the
mesh puts inside a directory is itself a declared resource, and orphans are removed in reverse
declaration order — so by the time a directory is reached, what the mesh wrote there is already
gone. Anything still present was put there by something else.
## Considered Options
1. **A `keep` flag on the directory.** A module declares which of its directories hold data, and
the host leaves those. **Rejected.** It is safe only when somebody remembered, and the failure
of forgetting is total and silent. A rule that protects data only when it was asked to is not
a rule about data, it is a rule about attentiveness — and this is the one place in the system
where being wrong does not recover.
2. **Never remove a directory.** Simple and unarguably safe. **Rejected**, narrowly: every module
ever assigned would leave its directories behind for ever, and a machine that accumulates
things nobody can account for is one where nobody can tell what is still in use. The clean-up
that is genuinely the mesh's is worth keeping.
3. **Remove a directory only when it is empty.** **Adopted.**
## Decision
**A directory that still holds anything is kept, and the mesh says so.** An empty one is removed.
**This is the host's existing line applied to the one shape where getting it wrong is
unrecoverable** — *it removes what it made and leaves what it merely configured*
([ADR 0005](0005-the-node-host.md)). An empty directory is what the host made. A full one is not.
**No flag, no declaration, nothing to remember.** Emptiness is the test, and it is derived from
the removal order rather than asserted: the mesh's own contents are gone by then, so what remains
is by definition something nobody declared.
**It is reported, not silent.** The outcome is `kept`, naming how many items are inside and saying
they are for a person to deal with. A directory quietly left behind is how a machine accumulates
things nobody can account for — which is the objection to option 2, and it is answered by saying
so rather than by deleting.
**Files are unchanged.** A declared file is the mesh's own — it wrote it, it owns it, and losing a
configuration file is not the failure this is about. The distinction is deliberate: **directories
hold what other things produced; files are what the mesh itself put there.**
## Consequences
**Moving a data directory is now safe by default.** The manifest changes, the old path stops being
declared, and the data stays where it is until somebody has looked at it. That was the operation
most likely to destroy something during the conversion, and it is now the operation that does the
least.
**Unassigning a module leaves its data.** Correct, and it means unassignment is no longer a way to
clean up — removing data is a person's act, done knowingly. Given what unassignment did before,
that is the trade being made and it is the right way round.
**A machine can accumulate directories nobody removed.** Accepted, and mitigated by saying so
every time rather than by a periodic sweep. A sweep would be the deletion this record exists to
prevent, on a timer, with nobody watching.
**It is not a backup, and must not be mistaken for one.** This stops the mesh destroying data. It
does nothing about a disk, a mistaken `rm`, or a service corrupting its own store. The conversion
still needs backups taken and **restored** before anything is moved — a backup nobody has restored
is a belief, not a copy.
## References
- [ADR 0005](0005-the-node-host.md) — the host removes what it made
- [`03-DESIGN/01-to-be/00-work-breakdown.md`](../03-DESIGN/01-to-be/00-work-breakdown.md) — the
conversion this was found by planning
+1
View File
@@ -94,6 +94,7 @@ python3 00-META/checks/index.py fail if stale
- **0008** — [A context owns its store, exclusively](0008-a-context-owns-its-store.md) - **0008** — [A context owns its store, exclusively](0008-a-context-owns-its-store.md)
- **0028** — [The substrate supplies the control plane and nothing else](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md) - **0028** — [The substrate supplies the control plane and nothing else](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md)
- **0029** — [A network is a shape, because an action cannot be undone](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md) - **0029** — [A network is a shape, because an action cannot be undone](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md)
- **0030** — [Data outlives the mesh that declared it](0030-data-outlives-the-mesh-that-declared-it.md)
### What runs on them, and how it gets there ### What runs on them, and how it gets there
+36 -2
View File
@@ -115,13 +115,47 @@ endpoint, so *a consumer cannot reach another consumer's bucket* is a policy som
a policy granting everything would pass every other test. **What is asserted is what the policy a policy granting everything would pass every other test. **What is asserted is what the policy
does not say.** does not say.**
## Data is the constraint, and it outranks the order below
*2026-08-31.* The modules being converted run live services — identity, mail — and **the data must
survive every step**. A data folder may move; it may never be lost.
**One thing was found by asking this and is fixed**
([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)): the host deleted
a directory and everything under it when the directory stopped being declared, which happens when
a module is unassigned or a manifest is edited to move a data folder — the exact operation this
plan needs. A directory holding anything the mesh did not put there is now kept and reported.
**That is not a backup and must not be read as one.** It stops the mesh destroying data. It does
nothing about a disk, a mistaken command, or a service corrupting its own store.
**So the rule for every step below:** the data is copied, the copy is verified by reading it back
through the service that owns it, and only then does anything point at the new location. Never
moved and then checked. **A backup nobody has restored is a belief, not a copy.**
## Where it starts, and what that costs
**On the node holding all the production data**, because that is where the services being
converted actually are.
Recorded plainly rather than argued with: this is the highest-risk order available. Everything
proven so far was proven on machines that could be destroyed and raised again, and the first real
exercise of the conversion will be on the one machine where a mistake is not recoverable. Nothing
about the lab work transfers automatically — a scenario proves the mechanism, not the state on
that machine.
**What makes it survivable is preparation rather than caution**: a restored backup before the
first step, one service at a time, and the previous arrangement left standing until the new one
has been read back. None of that is slower than the alternative, because the alternative includes
losing something.
## Phase 2 — the first real module ## Phase 2 — the first real module
| # | task | done when | | # | task | done when |
|---|---|---| |---|---|---|
| 2.1 | Port an **object store** module | it runs on the new mesh, serves a bucket to another module, and its credential rotates | | 2.1 | Port an **object store** module | it runs on the new mesh, serves a bucket to another module, and its credential rotates |
| 2.2 | Run it beside the existing one | both exist; nothing depends on the new one yet | | 2.2 | Copy the data, and read it back through the service that owns it | the new location answers with what the old one holds |
| 2.3 | Move one dependent onto it | something real reads and writes through the new mesh's copy | | 2.3 | Point one dependent at it, old arrangement left standing | something real reads and writes through the new mesh's copy |
**Checkpoint, and it is a human one:** it runs for a week before anything else moves. The point of **Checkpoint, and it is a human one:** it runs for a week before anything else moves. The point of
going first is to find what Phase 1 missed, and a week is roughly how long that takes to show. going first is to find what Phase 1 missed, and a week is roughly how long that takes to show.