Jochen: a normal application has 3-5 ADRs, maybe 10 for a large one, and we are at 65. Fair, and the cause is mine -- I recorded every FINDING as a decision rather than every fork in the road. Two merges, both cases where one decision had been split across many records because it was taken over several days rather than at once. 0019 absorbs ten records about how this repository works: what it is and that it is public, the folder flow, the two design layers, the issue front door, status in frontmatter, playbooks, the naming rule, the product name. Those were never ten decisions -- they were one, seen from ten angles as the repository took shape. 0016 absorbs the five about the lab: a node is a virtual machine, a router is scenery, a scenario declares the underlay, a scenario is a closed address space, and the two scenario classes. Same pattern -- one design, split by the order it was worked out in. The consolidated 0019 also raises the bar for what earns a record, since that is what produced 65: a record is warranted when there is a genuine fork -- a direction reversed, an alternative that will be proposed again, something contested. A finding is not a decision, and a bug is certainly not. Everything else belongs in the design document where the reasoning is actually read. The checker earned its place here. Deleting nine records left 13 dangling links across the repository and it named every one, including in AGENTS.md. Nothing was found by reading. Remaining clusters worth the same treatment: the host (8 records), delivery (5), modules (6), connectivity (4), substrate and control plane (4). That would be 52 down to roughly 30.
6.1 KiB
status, initiated, touches, became
| status | initiated | touches | became | ||||
|---|---|---|---|---|---|---|---|
| active | 2026-08-23 |
|
009 — Getting from the mesh that exists to the mesh that is designed
What is being investigated
How a running mesh becomes the one in research 006, without losing what it currently carries.
The proposed shape: build tiers 0, 1 and 2, then replace the current setup in one move.
Why big-bang is the right instinct here
Recorded because incremental is the reflex answer and it is wrong in this case.
- The two models are structurally incompatible. The tier rule, the host absorbing what are now modules, the artifact/part split, provisioning generalised to the control plane's own requirements — none of these can half-apply. Running both models at once means the old one's assumptions keep constraining the new one, which is how a migration becomes permanent.
- Nothing external depends on it. No users outside the operator, no service level to hold.
- The lab exists precisely for this (ADR 0016). A big-bang that has been rehearsed end to end, repeatedly, on identical machines is not the same risk as one performed for the first time on the real mesh. This is also why the lab is phase 0 rather than a verification step later: the new mesh is developed inside it, so by cutover the procedure has been run hundreds of times rather than rehearsed a few.
- Incremental would carry the rot forward. The as-is layer documents silent failure paths, a dead test harness and unenforced rules. A gradual migration preserves them by definition.
The distinction that lowers the risk
Replace the control plane; do not move the workloads.
The things that would hurt to lose — mail, media, source, databases, their data directories — are not the mesh. They are what the mesh manages. They sit in container volumes on nodes, and they do not need to move for the control plane above them to be replaced.
So the big-bang is: the old control plane stops managing these nodes, and the new one starts — with the workload data untouched, in place, and re-declared rather than migrated.
That reframing turns "replace the mesh" into "replace the part with no persistent state of its own", which is a materially smaller act than it first sounds.
The tension this exposes
Taking over already-running workloads is adoption, and adoption was ruled out of scope —
recorded as a legacy path in
03-DESIGN/01-to-be/01-end-to-end-testing.md.
The migration appears to need exactly the capability the design declared it would not have.
The way out, to be tested: the new mesh does not adopt anything. It declares the workloads from scratch and points them at data directories that already exist. Nothing inspects a running machine to learn what is there; the declarations are written from the as-is layer, which is what that layer is for. Data survives because it was never touched, not because it was adopted.
If that holds, adoption stays out of scope and the migration is ordinary declaration. If it does not, adoption needs a one-time, explicitly unsupported tool, and that should be a decision rather than a discovery.
Sequencing
| Phase | What | Done when |
|---|---|---|
| 0 | Build the lab's bootstrap scenario (ADR 0016) — virtual machines, a network, a way to place a binary, snapshot and reset. No forge, no coordinator, no pipeline. | A machine can be raised from nothing, reset, and raised again, repeatably. |
| A | Build tier 0, inside the lab. The host's interface first — it carries the skeleton's biggest unproven claim. | A bare machine becomes a managed node with no mesh present. |
| B | Build tier 1 and 2. The bootstrap scenario grows into the full one by addition — the same machines, with more placed inside them. | The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts. |
| C | Enough of tier 3 to operate it. | The mesh can be driven without direct database access. |
| D | Declare the existing workloads against the new model. | The lab runs them, with copies of real data shapes. |
| E | Rehearse the cutover in the lab against a mesh built to resemble the real one. | Repeatable, and repeatably reversible. |
| F | Cut over. | The real nodes are managed by the new control plane. |
| G | Reach self-hosting — re-bind delivery from external providers to the mesh's own forge and registries. | The mesh builds and deploys itself. |
Phase G is deliberately last. Per research 006, self-hosting is a state the mesh reaches; attempting the cutover and the self-hosting transition in the same move recreates exactly the circularity the skeleton removes — and would mean a failed cutover could take away the means to fix it.
The questions
| Question | Why it matters |
|---|---|
| What state must survive the cutover, versus be re-created? | Workload data must. Provisioned credentials could be re-issued. The mesh's own inventory could be re-declared. Each answer changes the risk. |
| What is the way back? | A cutover with no rollback is not a plan. If workload data is untouched, reverting may be as small as re-pointing the old control plane at it — to be verified, not assumed. |
| How is the cutover rehearsed against something resembling the real mesh, without copying the real mesh into a repository? | The lab must be able to model the real topology's shape without carrying its identity. |
| Does anything have to keep running during the cutover? | Mail and source are the obvious candidates. If yes, "big-bang" is really "big-bang with exceptions", and the exceptions should be named now. |
| Is the forge inside or outside the cutover? | If the mesh's own forge goes down with the old control plane, the means of deploying a fix goes with it. This is the self-hosting circularity appearing as a migration risk. |