Files
hq/01-RESEARCH/009-migration/00-overview.md
T
jschoubben b4904fec7e The lab comes first, and its first scenario has no pipeline
The lab was designed around a module under test, with a scenario being a
complete mesh — forge, coordinator, cascade, verify. That is unusable for
building the new mesh, because all four are tier 2 and do not exist yet.

And research 009 had the sequence backwards. It placed the lab at phase B
as verification of tiers already built, but tier 0 is the component that
takes over a machine's packages, services and network. It cannot be
developed against a machine anyone needs. The lab has to exist before the
thing it will test.

ADR 0029 splits scenarios into two classes. The bootstrap scenario is
virtual machines, the host binary and a pinned bundle, with the verdict
coming from what the host reports about the state it reconciled. The full
scenario is the designed one. The first is a strict subset of the second —
same virtualisation, same networking, same lifecycle, stopping before a
control plane exists — so the second is reached by addition rather than
rework.

The consequence worth having: raising a node from nothing stops being the
least-exercised path in the system and becomes the inner development loop.

It also settles the runner's two jobs. Scenario lifecycle is needed
immediately, because something must materialise and reset a mesh before
anything can be written against it. Assertion execution waits for the full
scenario.

Corrects a stale claim in the design while amending it: it argued
scenarios were affordable with system containers and would not be with
virtual machines. ADR 0016 superseded that reasoning and the text had not
followed.

Issue 007: the lab's first requirement is installed and unusable. The
virtualisation package is present and explicitly installed; both units are
disabled, the operator is in no group, and the client reports the server
unreachable. Not issue 001 again — that is an install failing while
reporting success. This is an install succeeding when success was not the
point. A package is files; a capability is a running service and an
identity permitted to reach it, and the module model has no vocabulary for
the second.
2026-08-23 21:57:14 +02:00

96 lines
6.1 KiB
Markdown

---
status: active
initiated: 2026-08-23
touches:
- 01-RESEARCH/006-mesh-from-scratch/code-skeleton.md
- 03-DESIGN/01-to-be/01-end-to-end-testing.md
- 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
- 03-DESIGN/00-as-is/00-overview.md
became: []
---
# 009 — Getting from the mesh that exists to the mesh that is designed
## What is being investigated
How a running mesh becomes the one in
[research 006](../006-mesh-from-scratch/code-skeleton.md), without losing what it currently
carries.
The proposed shape: **build tiers 0, 1 and 2, then replace the current setup in one move.**
## Why big-bang is the right instinct here
Recorded because incremental is the reflex answer and it is wrong in this case.
- **The two models are structurally incompatible.** The tier rule, the host absorbing what are
now modules, the artifact/part split, provisioning generalised to the control plane's own
requirements — none of these can half-apply. Running both models at once means the old one's
assumptions keep constraining the new one, which is how a migration becomes permanent.
- **Nothing external depends on it.** No users outside the operator, no service level to hold.
- **The lab exists precisely for this** ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)).
A big-bang that has been rehearsed end to end, repeatedly, on identical machines is not the
same risk as one performed for the first time on the real mesh. This is also why the lab is
phase 0 rather than a verification step later: the new mesh is *developed* inside it, so by
cutover the procedure has been run hundreds of times rather than rehearsed a few.
- **Incremental would carry the rot forward.** The as-is layer documents silent failure paths,
a dead test harness and unenforced rules. A gradual migration preserves them by definition.
## The distinction that lowers the risk
**Replace the control plane; do not move the workloads.**
The things that would hurt to lose — mail, media, source, databases, their data directories —
are not the mesh. They are what the mesh manages. They sit in container volumes on nodes, and
they do not need to move for the control plane above them to be replaced.
So the big-bang is: the old control plane stops managing these nodes, and the new one starts —
with the workload data untouched, in place, and re-declared rather than migrated.
That reframing turns "replace the mesh" into "replace the part with no persistent state of its
own", which is a materially smaller act than it first sounds.
## The tension this exposes
Taking over already-running workloads is **adoption**, and adoption was ruled out of scope —
recorded as a legacy path in
[`03-DESIGN/01-to-be/01-end-to-end-testing.md`](../../03-DESIGN/01-to-be/01-end-to-end-testing.md).
The migration appears to need exactly the capability the design declared it would not have.
The way out, to be tested: the new mesh does not adopt anything. It **declares** the workloads
from scratch and points them at data directories that already exist. Nothing inspects a running
machine to learn what is there; the declarations are written from the as-is layer, which is what
that layer is for. Data survives because it was never touched, not because it was adopted.
If that holds, adoption stays out of scope and the migration is ordinary declaration. If it does
not, adoption needs a one-time, explicitly unsupported tool, and that should be a decision rather
than a discovery.
## Sequencing
| Phase | What | Done when |
|---|---|---|
| **0** | **Build the lab's bootstrap scenario** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)) — virtual machines, a network, a way to place a binary, snapshot and reset. No forge, no coordinator, no pipeline. | A machine can be raised from nothing, reset, and raised again, repeatably. |
| A | Build tier 0, **inside the lab**. The host's interface first — it carries the skeleton's biggest unproven claim. | A bare machine becomes a managed node with no mesh present. |
| B | Build tier 1 and 2. The bootstrap scenario grows into the full one by addition — the same machines, with more placed inside them. | The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts. |
| C | Enough of tier 3 to operate it. | The mesh can be driven without direct database access. |
| D | Declare the existing workloads against the new model. | The lab runs them, with copies of real data shapes. |
| E | **Rehearse the cutover in the lab** against a mesh built to resemble the real one. | Repeatable, and repeatably reversible. |
| F | Cut over. | The real nodes are managed by the new control plane. |
| G | Reach self-hosting — re-bind delivery from external providers to the mesh's own forge and registries. | The mesh builds and deploys itself. |
Phase G is deliberately last. Per research 006, self-hosting is a state the mesh **reaches**;
attempting the cutover and the self-hosting transition in the same move recreates exactly the
circularity the skeleton removes — and would mean a failed cutover could take away the means to
fix it.
## The questions
| Question | Why it matters |
|---|---|
| What state must **survive** the cutover, versus be re-created? | Workload data must. Provisioned credentials could be re-issued. The mesh's own inventory could be re-declared. Each answer changes the risk. |
| What is the **way back**? | A cutover with no rollback is not a plan. If workload data is untouched, reverting may be as small as re-pointing the old control plane at it — to be verified, not assumed. |
| How is the cutover **rehearsed** against something resembling the real mesh, without copying the real mesh into a repository? | The lab must be able to model the real topology's shape without carrying its identity. |
| Does anything have to keep running **during** the cutover? | Mail and source are the obvious candidates. If yes, "big-bang" is really "big-bang with exceptions", and the exceptions should be named now. |
| Is the forge inside or outside the cutover? | If the mesh's own forge goes down with the old control plane, the means of deploying a fix goes with it. This is the self-hosting circularity appearing as a migration risk. |