Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24

Merged
jschoubben merged 177 commits from reconcile-init-into-main into main 2026-09-05 10:27:11 +00:00
Showing only changes of commit c570c687f6 - Show all commits
+37
View File
@@ -149,6 +149,43 @@ first step, one service at a time, and the previous arrangement left standing un
has been read back. None of that is slower than the alternative, because the alternative includes
losing something.
## How the two systems hand over
*2026-08-31.* **The old system's brain is switched off; its services keep running.**
Not a migration and not a period of dual control. The old control plane — provisioning, the
coordinator, the pipeline, the things that *decide* and *write* — is stopped. Every workload it
was managing goes on running exactly as it is, because nothing is managing it. Then the new mesh
takes ownership of them one at a time.
**Nothing is ever unassigned in the old system.** Unassigning is how it removes things, and
removing is how data is lost. The old system is never asked to take anything away; it is asked to
stop having opinions.
| | |
|---|---|
| **stopped, and disabled** | provisioning, the coordinator, environment and configuration sync, the pipeline — anything that decides or writes a file |
| **left alone entirely** | the units running the actual services: identity, mail, databases, the forge. They keep serving throughout |
| **never used** | unassign, remove, delete — any operation whose job is to take something away |
**Disabled, not merely stopped**, and this is the part that is easy to get wrong: those units are
enabled, so stopping them lasts until the machine reboots. A reboot mid-conversion would bring the
old control plane back and it would resume regenerating managed files underneath the new one —
which is the one situation where two systems really would be fighting over the same machine.
**A service left running with nothing managing it is the safe state.** It has its data, its
configuration is already on disk, and nothing is going to change either. That is the whole trick:
the risk in a conversion is in the *managing*, not in the *running*.
**A brief interruption is acceptable. Losing data is not.** Where those two trade against each
other, the interruption wins every time — a service can be restarted, and there is no operation
that un-deletes a mail spool.
**The new host cannot remove what it did not put there.** Orphans are per-origin, so it only ever
removes resources it recorded itself. Services it has never been told about are not orphans to
it — they are simply not its business, which is what makes taking ownership one module at a time
safe.
## Phase 2 — the first real module
| # | task | done when |