From c570c687f642b0f9a25f0a7c9fa6bc6349f8038b Mon Sep 17 00:00:00 2001 From: jochen Date: Mon, 31 Aug 2026 19:55:35 +0200 Subject: [PATCH] The handover: switch off the old brain, leave the services running MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The conversion method, recorded because it decides everything else and was not written down. The old control plane is stopped — provisioning, coordinator, syncs, the pipeline, anything that decides or writes. The workloads it was managing keep running, because nothing is managing them. The new mesh then takes ownership one module at a time. Nothing is ever unassigned in the old system. Unassigning is how it removes things and removing is how data is lost; it is asked to stop having opinions, never to take anything away. Disabled rather than merely stopped, which is the part easy to get wrong: those units are enabled, so a stop lasts until the next reboot. A reboot mid-conversion would bring the old control plane back to regenerate managed files underneath the new one — the one situation where two systems really would fight over a machine. A service left running with nothing managing it is the safe state: it has its data, its configuration is on disk, and nothing will change either. The risk in a conversion is in the managing, not the running. Also records why taking ownership piecemeal is safe: the new host's orphan removal is per-origin, so it only removes what it recorded itself. Services it was never told about are not orphans to it. --- 03-DESIGN/01-to-be/00-work-breakdown.md | 37 +++++++++++++++++++++++++ 1 file changed, 37 insertions(+) diff --git a/03-DESIGN/01-to-be/00-work-breakdown.md b/03-DESIGN/01-to-be/00-work-breakdown.md index e458d96..cf61b0b 100644 --- a/03-DESIGN/01-to-be/00-work-breakdown.md +++ b/03-DESIGN/01-to-be/00-work-breakdown.md @@ -149,6 +149,43 @@ first step, one service at a time, and the previous arrangement left standing un has been read back. None of that is slower than the alternative, because the alternative includes losing something. +## How the two systems hand over + +*2026-08-31.* **The old system's brain is switched off; its services keep running.** + +Not a migration and not a period of dual control. The old control plane — provisioning, the +coordinator, the pipeline, the things that *decide* and *write* — is stopped. Every workload it +was managing goes on running exactly as it is, because nothing is managing it. Then the new mesh +takes ownership of them one at a time. + +**Nothing is ever unassigned in the old system.** Unassigning is how it removes things, and +removing is how data is lost. The old system is never asked to take anything away; it is asked to +stop having opinions. + +| | | +|---|---| +| **stopped, and disabled** | provisioning, the coordinator, environment and configuration sync, the pipeline — anything that decides or writes a file | +| **left alone entirely** | the units running the actual services: identity, mail, databases, the forge. They keep serving throughout | +| **never used** | unassign, remove, delete — any operation whose job is to take something away | + +**Disabled, not merely stopped**, and this is the part that is easy to get wrong: those units are +enabled, so stopping them lasts until the machine reboots. A reboot mid-conversion would bring the +old control plane back and it would resume regenerating managed files underneath the new one — +which is the one situation where two systems really would be fighting over the same machine. + +**A service left running with nothing managing it is the safe state.** It has its data, its +configuration is already on disk, and nothing is going to change either. That is the whole trick: +the risk in a conversion is in the *managing*, not in the *running*. + +**A brief interruption is acceptable. Losing data is not.** Where those two trade against each +other, the interruption wins every time — a service can be restarted, and there is no operation +that un-deletes a mail spool. + +**The new host cannot remove what it did not put there.** Orphans are per-origin, so it only ever +removes resources it recorded itself. Services it has never been told about are not orphans to +it — they are simply not its business, which is what makes taking ownership one module at a time +safe. + ## Phase 2 — the first real module | # | task | done when |