Files
hq/01-RESEARCH/009-migration/00-overview.md
T
jschoubben e1febe8e0f Renumber the records 1 to 23
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18,
19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only
the archaeology of what used to be there.

Renumbered contiguously. Renames run in ascending order, so every target number
is already free and no two files ever collide.

The reference rewrite is one simultaneous pass rather than a sequence of
replacements. Numbers moved into slots other numbers were vacating -- the node
host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time
would have cascaded and silently pointed things at the wrong record.

Seven plain-text references survived the merges as prose rather than links,
naming records that no longer existed: the enrolment token, the link boundary,
what a declaration is, reachability, the repository structure. Each mapped to
the consolidated record that now holds it.

Verified rather than assumed: every [ADR NNNN](path) link now has matching text
and target, checked across the whole repository, and the checker passes.

Frontmatter `consolidates:` lists dropped -- they named records that are gone,
and each consolidated record already says in prose what it absorbed.
2026-08-28 23:28:34 +02:00

6.1 KiB

status, initiated, touches, became
status initiated touches became
active 2026-08-23
01-RESEARCH/006-mesh-from-scratch/code-skeleton.md
03-DESIGN/01-to-be/01-end-to-end-testing.md
02-DECISIONS/0009-the-lab.md
03-DESIGN/00-as-is/00-overview.md

009 — Getting from the mesh that exists to the mesh that is designed

What is being investigated

How a running mesh becomes the one in research 006, without losing what it currently carries.

The proposed shape: build tiers 0, 1 and 2, then replace the current setup in one move.

Why big-bang is the right instinct here

Recorded because incremental is the reflex answer and it is wrong in this case.

  • The two models are structurally incompatible. The tier rule, the host absorbing what are now modules, the artifact/part split, provisioning generalised to the control plane's own requirements — none of these can half-apply. Running both models at once means the old one's assumptions keep constraining the new one, which is how a migration becomes permanent.
  • Nothing external depends on it. No users outside the operator, no service level to hold.
  • The lab exists precisely for this (ADR 0009). A big-bang that has been rehearsed end to end, repeatedly, on identical machines is not the same risk as one performed for the first time on the real mesh. This is also why the lab is phase 0 rather than a verification step later: the new mesh is developed inside it, so by cutover the procedure has been run hundreds of times rather than rehearsed a few.
  • Incremental would carry the rot forward. The as-is layer documents silent failure paths, a dead test harness and unenforced rules. A gradual migration preserves them by definition.

The distinction that lowers the risk

Replace the control plane; do not move the workloads.

The things that would hurt to lose — mail, media, source, databases, their data directories — are not the mesh. They are what the mesh manages. They sit in container volumes on nodes, and they do not need to move for the control plane above them to be replaced.

So the big-bang is: the old control plane stops managing these nodes, and the new one starts — with the workload data untouched, in place, and re-declared rather than migrated.

That reframing turns "replace the mesh" into "replace the part with no persistent state of its own", which is a materially smaller act than it first sounds.

The tension this exposes

Taking over already-running workloads is adoption, and adoption was ruled out of scope — recorded as a legacy path in 03-DESIGN/01-to-be/01-end-to-end-testing.md. The migration appears to need exactly the capability the design declared it would not have.

The way out, to be tested: the new mesh does not adopt anything. It declares the workloads from scratch and points them at data directories that already exist. Nothing inspects a running machine to learn what is there; the declarations are written from the as-is layer, which is what that layer is for. Data survives because it was never touched, not because it was adopted.

If that holds, adoption stays out of scope and the migration is ordinary declaration. If it does not, adoption needs a one-time, explicitly unsupported tool, and that should be a decision rather than a discovery.

Sequencing

Phase What Done when
0 Build the lab's bootstrap scenario (ADR 0009) — virtual machines, a network, a way to place a binary, snapshot and reset. No forge, no coordinator, no pipeline. A machine can be raised from nothing, reset, and raised again, repeatably.
A Build tier 0, inside the lab. The host's interface first — it carries the skeleton's biggest unproven claim. A bare machine becomes a managed node with no mesh present.
B Build tier 1 and 2. The bootstrap scenario grows into the full one by addition — the same machines, with more placed inside them. The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts.
C Enough of tier 3 to operate it. The mesh can be driven without direct database access.
D Declare the existing workloads against the new model. The lab runs them, with copies of real data shapes.
E Rehearse the cutover in the lab against a mesh built to resemble the real one. Repeatable, and repeatably reversible.
F Cut over. The real nodes are managed by the new control plane.
G Reach self-hosting — re-bind delivery from external providers to the mesh's own forge and registries. The mesh builds and deploys itself.

Phase G is deliberately last. Per research 006, self-hosting is a state the mesh reaches; attempting the cutover and the self-hosting transition in the same move recreates exactly the circularity the skeleton removes — and would mean a failed cutover could take away the means to fix it.

The questions

Question Why it matters
What state must survive the cutover, versus be re-created? Workload data must. Provisioned credentials could be re-issued. The mesh's own inventory could be re-declared. Each answer changes the risk.
What is the way back? A cutover with no rollback is not a plan. If workload data is untouched, reverting may be as small as re-pointing the old control plane at it — to be verified, not assumed.
How is the cutover rehearsed against something resembling the real mesh, without copying the real mesh into a repository? The lab must be able to model the real topology's shape without carrying its identity.
Does anything have to keep running during the cutover? Mail and source are the obvious candidates. If yes, "big-bang" is really "big-bang with exceptions", and the exceptions should be named now.
Is the forge inside or outside the cutover? If the mesh's own forge goes down with the old control plane, the means of deploying a fix goes with it. This is the self-hosting circularity appearing as a migration risk.