Files
hq/01-RESEARCH/009-migration/00-overview.md
T
jschoubben b4904fec7e The lab comes first, and its first scenario has no pipeline
The lab was designed around a module under test, with a scenario being a
complete mesh — forge, coordinator, cascade, verify. That is unusable for
building the new mesh, because all four are tier 2 and do not exist yet.

And research 009 had the sequence backwards. It placed the lab at phase B
as verification of tiers already built, but tier 0 is the component that
takes over a machine's packages, services and network. It cannot be
developed against a machine anyone needs. The lab has to exist before the
thing it will test.

ADR 0029 splits scenarios into two classes. The bootstrap scenario is
virtual machines, the host binary and a pinned bundle, with the verdict
coming from what the host reports about the state it reconciled. The full
scenario is the designed one. The first is a strict subset of the second —
same virtualisation, same networking, same lifecycle, stopping before a
control plane exists — so the second is reached by addition rather than
rework.

The consequence worth having: raising a node from nothing stops being the
least-exercised path in the system and becomes the inner development loop.

It also settles the runner's two jobs. Scenario lifecycle is needed
immediately, because something must materialise and reset a mesh before
anything can be written against it. Assertion execution waits for the full
scenario.

Corrects a stale claim in the design while amending it: it argued
scenarios were affordable with system containers and would not be with
virtual machines. ADR 0016 superseded that reasoning and the text had not
followed.

Issue 007: the lab's first requirement is installed and unusable. The
virtualisation package is present and explicitly installed; both units are
disabled, the operator is in no group, and the client reports the server
unreachable. Not issue 001 again — that is an install failing while
reporting success. This is an install succeeding when success was not the
point. A package is files; a capability is a running service and an
identity permitted to reach it, and the module model has no vocabulary for
the second.
2026-08-23 21:57:14 +02:00

6.1 KiB

status, initiated, touches, became
status initiated touches became
active 2026-08-23
01-RESEARCH/006-mesh-from-scratch/code-skeleton.md
03-DESIGN/01-to-be/01-end-to-end-testing.md
02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
03-DESIGN/00-as-is/00-overview.md

009 — Getting from the mesh that exists to the mesh that is designed

What is being investigated

How a running mesh becomes the one in research 006, without losing what it currently carries.

The proposed shape: build tiers 0, 1 and 2, then replace the current setup in one move.

Why big-bang is the right instinct here

Recorded because incremental is the reflex answer and it is wrong in this case.

  • The two models are structurally incompatible. The tier rule, the host absorbing what are now modules, the artifact/part split, provisioning generalised to the control plane's own requirements — none of these can half-apply. Running both models at once means the old one's assumptions keep constraining the new one, which is how a migration becomes permanent.
  • Nothing external depends on it. No users outside the operator, no service level to hold.
  • The lab exists precisely for this (ADR 0016). A big-bang that has been rehearsed end to end, repeatedly, on identical machines is not the same risk as one performed for the first time on the real mesh. This is also why the lab is phase 0 rather than a verification step later: the new mesh is developed inside it, so by cutover the procedure has been run hundreds of times rather than rehearsed a few.
  • Incremental would carry the rot forward. The as-is layer documents silent failure paths, a dead test harness and unenforced rules. A gradual migration preserves them by definition.

The distinction that lowers the risk

Replace the control plane; do not move the workloads.

The things that would hurt to lose — mail, media, source, databases, their data directories — are not the mesh. They are what the mesh manages. They sit in container volumes on nodes, and they do not need to move for the control plane above them to be replaced.

So the big-bang is: the old control plane stops managing these nodes, and the new one starts — with the workload data untouched, in place, and re-declared rather than migrated.

That reframing turns "replace the mesh" into "replace the part with no persistent state of its own", which is a materially smaller act than it first sounds.

The tension this exposes

Taking over already-running workloads is adoption, and adoption was ruled out of scope — recorded as a legacy path in 03-DESIGN/01-to-be/01-end-to-end-testing.md. The migration appears to need exactly the capability the design declared it would not have.

The way out, to be tested: the new mesh does not adopt anything. It declares the workloads from scratch and points them at data directories that already exist. Nothing inspects a running machine to learn what is there; the declarations are written from the as-is layer, which is what that layer is for. Data survives because it was never touched, not because it was adopted.

If that holds, adoption stays out of scope and the migration is ordinary declaration. If it does not, adoption needs a one-time, explicitly unsupported tool, and that should be a decision rather than a discovery.

Sequencing

Phase What Done when
0 Build the lab's bootstrap scenario (ADR 0029) — virtual machines, a network, a way to place a binary, snapshot and reset. No forge, no coordinator, no pipeline. A machine can be raised from nothing, reset, and raised again, repeatably.
A Build tier 0, inside the lab. The host's interface first — it carries the skeleton's biggest unproven claim. A bare machine becomes a managed node with no mesh present.
B Build tier 1 and 2. The bootstrap scenario grows into the full one by addition — the same machines, with more placed inside them. The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts.
C Enough of tier 3 to operate it. The mesh can be driven without direct database access.
D Declare the existing workloads against the new model. The lab runs them, with copies of real data shapes.
E Rehearse the cutover in the lab against a mesh built to resemble the real one. Repeatable, and repeatably reversible.
F Cut over. The real nodes are managed by the new control plane.
G Reach self-hosting — re-bind delivery from external providers to the mesh's own forge and registries. The mesh builds and deploys itself.

Phase G is deliberately last. Per research 006, self-hosting is a state the mesh reaches; attempting the cutover and the self-hosting transition in the same move recreates exactly the circularity the skeleton removes — and would mean a failed cutover could take away the means to fix it.

The questions

Question Why it matters
What state must survive the cutover, versus be re-created? Workload data must. Provisioned credentials could be re-issued. The mesh's own inventory could be re-declared. Each answer changes the risk.
What is the way back? A cutover with no rollback is not a plan. If workload data is untouched, reverting may be as small as re-pointing the old control plane at it — to be verified, not assumed.
How is the cutover rehearsed against something resembling the real mesh, without copying the real mesh into a repository? The lab must be able to model the real topology's shape without carrying its identity.
Does anything have to keep running during the cutover? Mail and source are the obvious candidates. If yes, "big-bang" is really "big-bang with exceptions", and the exceptions should be named now.
Is the forge inside or outside the cutover? If the mesh's own forge goes down with the old control plane, the means of deploying a fix goes with it. This is the self-hosting circularity appearing as a migration risk.