ADR 0027 — the product is Novox Mesh, shortened to mesh internally. HAL was never chosen: it arrived with the dotfiles repository this grew out of, it is borrowed, and it is borrowed from the canonical untrustworthy machine intelligence, which is an odd flag for infrastructure trusted with credentials. Timing is the substance of the decision, not an aside — the skeleton is not built, so renaming costs a search and replace now and a migration later. Nox is an identity of Novox, and specifically the agent of the MESH rather than of a node. Nodes keep their own identities. Nox addresses them, and a human mostly talks to Nox — which makes it the concrete form of the mission's vision: state an intent, and the mesh works out which node holds the thing. It holds no private channel. The gap this opens is recorded: ADR 0012 binds every agent to a home node, and a mesh-scoped agent has none, so the model needs extending. ADR 0028 — HQ is company-scoped, novox/hq, with the mesh as its first product. Checked rather than assumed: the company organisation already holds live projects that the mesh builds and deploys, so they are tenants rather than peers, and the mesh is the ground they stand on. There is also company work outside the mesh already, which strengthens the case and means the eventual split is closer than "some day" — so each document's scope is fixed now, in a table, making that split mechanical instead of archaeological. The folders are deliberately not restructured yet. The skeleton takes the new vocabulary: mesh-host, mesh-substrate, mesh-control, mesh-surfaces, mesh-catalog. Substrate drops to four services now that identity is a hosted workload rather than a dependency. Research 009 opens the migration, with the reframing that lowers its risk: replace the control plane, do not move the workloads. Their data never moves, so it is re-declared rather than adopted — which keeps adoption out of scope, as the lab design requires. Self-hosting is the last phase, or a failed cutover takes away the means to fix it.
5.5 KiB
status, initiated, touches, became
| status | initiated | touches | became | ||||
|---|---|---|---|---|---|---|---|
| active | 2026-08-23 |
|
009 — Getting from the mesh that exists to the mesh that is designed
What is being investigated
How a running mesh becomes the one in research 006, without losing what it currently carries.
The proposed shape: build tiers 0, 1 and 2, then replace the current setup in one move.
Why big-bang is the right instinct here
Recorded because incremental is the reflex answer and it is wrong in this case.
- The two models are structurally incompatible. The tier rule, the host absorbing what are now modules, the artifact/part split, provisioning generalised to the control plane's own requirements — none of these can half-apply. Running both models at once means the old one's assumptions keep constraining the new one, which is how a migration becomes permanent.
- Nothing external depends on it. No users outside the operator, no service level to hold.
- The lab exists precisely for this (ADR 0016). A big-bang that has been rehearsed end to end, repeatedly, on identical machines is not the same risk as one performed for the first time on the real mesh.
- Incremental would carry the rot forward. The as-is layer documents silent failure paths, a dead test harness and unenforced rules. A gradual migration preserves them by definition.
The distinction that lowers the risk
Replace the control plane; do not move the workloads.
The things that would hurt to lose — mail, media, source, databases, their data directories — are not the mesh. They are what the mesh manages. They sit in container volumes on nodes, and they do not need to move for the control plane above them to be replaced.
So the big-bang is: the old control plane stops managing these nodes, and the new one starts — with the workload data untouched, in place, and re-declared rather than migrated.
That reframing turns "replace the mesh" into "replace the part with no persistent state of its own", which is a materially smaller act than it first sounds.
The tension this exposes
Taking over already-running workloads is adoption, and adoption was ruled out of scope —
recorded as a legacy path in
03-DESIGN/01-to-be/01-end-to-end-testing.md.
The migration appears to need exactly the capability the design declared it would not have.
The way out, to be tested: the new mesh does not adopt anything. It declares the workloads from scratch and points them at data directories that already exist. Nothing inspects a running machine to learn what is there; the declarations are written from the as-is layer, which is what that layer is for. Data survives because it was never touched, not because it was adopted.
If that holds, adoption stays out of scope and the migration is ordinary declaration. If it does not, adoption needs a one-time, explicitly unsupported tool, and that should be a decision rather than a discovery.
Sequencing
| Phase | What | Done when |
|---|---|---|
| A | Build tier 0. The host's interface first — it carries the skeleton's biggest unproven claim. | A bare machine becomes a managed node with no mesh present. |
| B | Build tier 1 and 2. | The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts. |
| C | Enough of tier 3 to operate it. | The mesh can be driven without direct database access. |
| D | Declare the existing workloads against the new model. | The lab runs them, with copies of real data shapes. |
| E | Rehearse the cutover in the lab against a mesh built to resemble the real one. | Repeatable, and repeatably reversible. |
| F | Cut over. | The real nodes are managed by the new control plane. |
| G | Reach self-hosting — re-bind delivery from external providers to the mesh's own forge and registries. | The mesh builds and deploys itself. |
Phase G is deliberately last. Per research 006, self-hosting is a state the mesh reaches; attempting the cutover and the self-hosting transition in the same move recreates exactly the circularity the skeleton removes — and would mean a failed cutover could take away the means to fix it.
The questions
| Question | Why it matters |
|---|---|
| What state must survive the cutover, versus be re-created? | Workload data must. Provisioned credentials could be re-issued. The mesh's own inventory could be re-declared. Each answer changes the risk. |
| What is the way back? | A cutover with no rollback is not a plan. If workload data is untouched, reverting may be as small as re-pointing the old control plane at it — to be verified, not assumed. |
| How is the cutover rehearsed against something resembling the real mesh, without copying the real mesh into a repository? | The lab must be able to model the real topology's shape without carrying its identity. |
| Does anything have to keep running during the cutover? | Mail and source are the obvious candidates. If yes, "big-bang" is really "big-bang with exceptions", and the exceptions should be named now. |
| Is the forge inside or outside the cutover? | If the mesh's own forge goes down with the old control plane, the means of deploying a fix goes with it. This is the self-hosting circularity appearing as a migration risk. |