Files
hq/01-RESEARCH/009-migration/00-overview.md
T
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00

6.1 KiB

status, initiated, touches, became
status initiated touches became
active 2026-08-23
01-RESEARCH/006-mesh-from-scratch/code-skeleton.md
03-DESIGN/01-to-be/01-end-to-end-testing.md
02-DECISIONS/0016-the-lab.md
03-DESIGN/00-as-is/00-overview.md

009 — Getting from the mesh that exists to the mesh that is designed

What is being investigated

How a running mesh becomes the one in research 006, without losing what it currently carries.

The proposed shape: build tiers 0, 1 and 2, then replace the current setup in one move.

Why big-bang is the right instinct here

Recorded because incremental is the reflex answer and it is wrong in this case.

  • The two models are structurally incompatible. The tier rule, the host absorbing what are now modules, the artifact/part split, provisioning generalised to the control plane's own requirements — none of these can half-apply. Running both models at once means the old one's assumptions keep constraining the new one, which is how a migration becomes permanent.
  • Nothing external depends on it. No users outside the operator, no service level to hold.
  • The lab exists precisely for this (ADR 0016). A big-bang that has been rehearsed end to end, repeatedly, on identical machines is not the same risk as one performed for the first time on the real mesh. This is also why the lab is phase 0 rather than a verification step later: the new mesh is developed inside it, so by cutover the procedure has been run hundreds of times rather than rehearsed a few.
  • Incremental would carry the rot forward. The as-is layer documents silent failure paths, a dead test harness and unenforced rules. A gradual migration preserves them by definition.

The distinction that lowers the risk

Replace the control plane; do not move the workloads.

The things that would hurt to lose — mail, media, source, databases, their data directories — are not the mesh. They are what the mesh manages. They sit in container volumes on nodes, and they do not need to move for the control plane above them to be replaced.

So the big-bang is: the old control plane stops managing these nodes, and the new one starts — with the workload data untouched, in place, and re-declared rather than migrated.

That reframing turns "replace the mesh" into "replace the part with no persistent state of its own", which is a materially smaller act than it first sounds.

The tension this exposes

Taking over already-running workloads is adoption, and adoption was ruled out of scope — recorded as a legacy path in 03-DESIGN/01-to-be/01-end-to-end-testing.md. The migration appears to need exactly the capability the design declared it would not have.

The way out, to be tested: the new mesh does not adopt anything. It declares the workloads from scratch and points them at data directories that already exist. Nothing inspects a running machine to learn what is there; the declarations are written from the as-is layer, which is what that layer is for. Data survives because it was never touched, not because it was adopted.

If that holds, adoption stays out of scope and the migration is ordinary declaration. If it does not, adoption needs a one-time, explicitly unsupported tool, and that should be a decision rather than a discovery.

Sequencing

Phase What Done when
0 Build the lab's bootstrap scenario (ADR 0016) — virtual machines, a network, a way to place a binary, snapshot and reset. No forge, no coordinator, no pipeline. A machine can be raised from nothing, reset, and raised again, repeatably.
A Build tier 0, inside the lab. The host's interface first — it carries the skeleton's biggest unproven claim. A bare machine becomes a managed node with no mesh present.
B Build tier 1 and 2. The bootstrap scenario grows into the full one by addition — the same machines, with more placed inside them. The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts.
C Enough of tier 3 to operate it. The mesh can be driven without direct database access.
D Declare the existing workloads against the new model. The lab runs them, with copies of real data shapes.
E Rehearse the cutover in the lab against a mesh built to resemble the real one. Repeatable, and repeatably reversible.
F Cut over. The real nodes are managed by the new control plane.
G Reach self-hosting — re-bind delivery from external providers to the mesh's own forge and registries. The mesh builds and deploys itself.

Phase G is deliberately last. Per research 006, self-hosting is a state the mesh reaches; attempting the cutover and the self-hosting transition in the same move recreates exactly the circularity the skeleton removes — and would mean a failed cutover could take away the means to fix it.

The questions

Question Why it matters
What state must survive the cutover, versus be re-created? Workload data must. Provisioned credentials could be re-issued. The mesh's own inventory could be re-declared. Each answer changes the risk.
What is the way back? A cutover with no rollback is not a plan. If workload data is untouched, reverting may be as small as re-pointing the old control plane at it — to be verified, not assumed.
How is the cutover rehearsed against something resembling the real mesh, without copying the real mesh into a repository? The lab must be able to model the real topology's shape without carrying its identity.
Does anything have to keep running during the cutover? Mail and source are the obvious candidates. If yes, "big-bang" is really "big-bang with exceptions", and the exceptions should be named now.
Is the forge inside or outside the cutover? If the mesh's own forge goes down with the old control plane, the means of deploying a fix goes with it. This is the self-hosting circularity appearing as a migration risk.