"control plane" -> controller and "substrate" -> foundation throughout 03-DESIGN, 00-META and the README, with 06-the-control-plane.md and 07-the-substrate.md renamed to 06-the-controller.md and 07-the-foundation.md. The immutable 02-DECISIONS records keep their original wording (and links to them are unchanged) — a term retired here may still appear there, which the glossary explains how to read. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
431 lines
25 KiB
Markdown
431 lines
25 KiB
Markdown
---
|
||
layer: to-be
|
||
status: designed
|
||
code: []
|
||
updated: 2026-09-01
|
||
decisions:
|
||
- 02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md
|
||
- 02-DECISIONS/0005-the-node-host.md
|
||
- 02-DECISIONS/0009-modules-and-the-graph.md
|
||
- 02-DECISIONS/0016-the-lab.md
|
||
- 02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md
|
||
---
|
||
|
||
# Work breakdown — replacing what provisions the mesh
|
||
|
||
*Rewritten 2026-08-31. The previous version planned a decomposition of the existing system in
|
||
place: extract contexts, convert modules to declared features, shrink its shared library. That is
|
||
not what is being done — a replacement is being built beside it, and the old plan's Phase 0 was
|
||
the only part that survived contact with it. So the document that was supposed to say what happens
|
||
next had been describing work on a system being retired.*
|
||
|
||
## The goal, in one sentence
|
||
|
||
**Modules move to the new mesh one at a time, until the old registry can be switched off.**
|
||
|
||
Everything below is ordered by what that requires. Nothing here is a rewrite of the old system;
|
||
its modules are the input.
|
||
|
||
## Phase 0 — a mesh that runs — **done**
|
||
|
||
Not *the code exists*. Twenty-two assertions on real machines in the lab, each confirmed to fail
|
||
when the behaviour is removed ([ADR 0016](../../02-DECISIONS/0016-the-lab.md),
|
||
[ADR 0017](../../02-DECISIONS/0017-a-test-defends-a-decision.md)).
|
||
|
||
| what is proven | |
|
||
|---|---|
|
||
| **a mesh comes into being** | a bare machine becomes one; others join with nothing but a token |
|
||
| **credentials** | delivered to both ends with the mesh holding neither; rotated so the old one stops working |
|
||
| **declarations survive reality** | a stopped machine is waited for; one that fell behind catches up unnamed; unassigning takes away exactly what it should; what the mesh says nothing about is left alone |
|
||
| **failure is legible** | a machine that cannot do what it was told is named, with why |
|
||
| **the mesh runs itself** | its own artifact store, and a builder that is a module the mesh assigns |
|
||
| **names and reachability** | internal names, wildcards under a machine, containers reaching other machines, certificates the mesh issued, filtering that matches exactly what was declared |
|
||
| **delivery** | a new commit reaches a machine already running the old one |
|
||
| **model access** | answered by a record, with a key the mesh cannot read |
|
||
|
||
**What Phase 0 does not prove, and it is the important sentence in this document:** every module
|
||
exercised above was written to test the mechanism. **No module from the existing system has ever
|
||
run on this.** The vocabulary was shaped by the things used to test it — the same fault as a
|
||
fixture agreeing with the code it checks
|
||
([`04-ISSUES/005`](../../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)), at the
|
||
scale of a design.
|
||
|
||
## Phase 1 — the vocabulary a real module needs
|
||
|
||
Found by taking real modules and asking what they would require. Each is a gap in what can be
|
||
*expressed*, not a defect in what is built.
|
||
|
||
| # | task | done when |
|
||
|---|---|---|
|
||
| ~~1.1~~ | ~~An **object-store provision**~~ — **done 2026-08-31**, and it needed no change to the mesh: see below | seven assertions against a real store |
|
||
| ~~1.2~~ | ~~**A session as a consumer of a licence**~~ — **done 2026-08-31**, and it also needed no change: see below | two sessions on one machine, different licences, each its own key |
|
||
| ~~1.3~~ | ~~A **network** shape, and ordering within a module~~ — **done 2026-08-31** | the shape is created and removed; ordering was already there, and is now asserted |
|
||
| 1.4 | **Public certificate issuance** — **built; one gap** | ordering, the challenge and issuance are proven against a real authority; **collecting the issued certificate is not** ([`04-ISSUES/020`](../../04-ISSUES/020-a-certificate-is-issued-and-never-collected/00-report.md)) |
|
||
|
||
**1.3 and 1.4 block later ones** and are listed now so they are not met as surprises. 1.3 is what
|
||
a mail system needs and nothing else so far does.
|
||
|
||
**Phase 1 is closed with 1.4 partly open**, deliberately. Two of its four items needed no code at
|
||
all; the network shape was built; and certificates are configured correctly, order correctly, and
|
||
are issued correctly — the client does not collect what the authority issued, against a server
|
||
that exists to be a test server. That is filed rather than chased, because the remainder may say
|
||
nothing about a real authority and the next thing to learn comes from moving a module rather than
|
||
from a fourth lab run.
|
||
|
||
**Checkpoint:** each is demonstrated in the lab before the module needing it is attempted.
|
||
|
||
### 1.2, and the same surprise twice
|
||
|
||
**A binding is per module per machine, and the two sessions are two modules** — the same mechanism
|
||
in different context roots, and a context root is what a module delivers. So `(node, module)`
|
||
already names them apart, and nothing needed adding.
|
||
[`14-model-access.md`](14-model-access.md) had called per-module-per-machine *a step toward it and
|
||
not it*, which is true of a **worker** — many run on one machine from one module — and not true of
|
||
a session, of which there is one per node and one for the mesh.
|
||
|
||
### 1.3, and the first one that needed building
|
||
|
||
**Ordering was already there** — the apply loop sorts nothing, so a module says *this before that*
|
||
by writing it first. Untested until now, and the kind of property a later change breaks silently.
|
||
Worth separating from readiness: a container started is not a container ready, and nothing waits.
|
||
What needs something *usable* retries, which is what both provisioners do and is the better answer
|
||
anyway, because a dependency can restart long after everything was applied.
|
||
|
||
**The network was a real gap, and the first thing in Phase 1 that needed a decision.** Adding a
|
||
shape widens what a compromised controller can express, so
|
||
[ADR 0029](../../02-DECISIONS/0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md)
|
||
records why this one is worth it: an `action` could create a network and **nothing could ever
|
||
remove it**, because an action leaves no footprint the host can undo. The vocabulary is nine.
|
||
|
||
**Three tasks in a row that were already possible.** Both were written from the design rather than
|
||
from the code, which is the review's finding arriving in the plan: *a claim here is counted, not
|
||
reasoned.* The remaining Phase 1 items should be checked against the code before being started,
|
||
not after.
|
||
|
||
### 1.1, and what it turned out to be
|
||
|
||
*Done 2026-08-31. Worth recording because the task was not the one written down.*
|
||
|
||
**The controller special-cases nothing.** `provides`, `requires`, `contributes` and `grants`
|
||
are name-agnostic — asking for a bucket needed no change to the mesh at all. What was missing was
|
||
a provider, and the last step where something on the machine turns a delivered secret into a key
|
||
that works. So "add an object-store provision" was never mesh work.
|
||
|
||
The provision is `s3-bucket`: a consumer's code is written against the S3 API and swapping one
|
||
store for another does not break it, so by
|
||
[ADR 0027](../../02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md) the name
|
||
says the protocol. A database is the other case, and names the engine.
|
||
|
||
**One assertion here that a database does not need.** One PostgreSQL server holds separate
|
||
databases and the product enforces the boundary; one object store holds every bucket behind one
|
||
endpoint, so *a consumer cannot reach another consumer's bucket* is a policy somebody wrote — and
|
||
a policy granting everything would pass every other test. **What is asserted is what the policy
|
||
does not say.**
|
||
|
||
## Data is the constraint, and it outranks the order below
|
||
|
||
*2026-08-31.* The modules being converted run live services — identity, mail — and **the data must
|
||
survive every step**. A data folder may move; it may never be lost.
|
||
|
||
**One thing was found by asking this and is fixed**
|
||
([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)): the host deleted
|
||
a directory and everything under it when the directory stopped being declared, which happens when
|
||
a module is unassigned or a manifest is edited to move a data folder — the exact operation this
|
||
plan needs. A directory holding anything the mesh did not put there is now kept and reported.
|
||
|
||
**That is not a backup and must not be read as one.** It stops the mesh destroying data. It does
|
||
nothing about a disk, a mistaken command, or a service corrupting its own store.
|
||
|
||
**So the rule for every step below:** the data is copied, the copy is verified by reading it back
|
||
through the service that owns it, and only then does anything point at the new location. Never
|
||
moved and then checked. **A backup nobody has restored is a belief, not a copy.**
|
||
|
||
## A module is adopted with the credentials it already has
|
||
|
||
*2026-08-31.* **Nothing is rotated during the conversion.** A service being adopted keeps the
|
||
password it is already using, because minting a new one is how a running service stops being able
|
||
to reach its own database in the middle of a migration.
|
||
|
||
The mesh has both paths and this needs the second:
|
||
|
||
| | |
|
||
|---|---|
|
||
| **generate** | a new secret, sealed to both ends. What a *new* module gets |
|
||
| **accept** | a value supplied from outside, sealed, plaintext discarded. **What an adopted module gets** |
|
||
|
||
**Rotation is a separate act, afterwards, once everything works.** The machinery for it is built
|
||
and proven — a credential moving at both ends with the old one ceasing to work — and it is exactly
|
||
the sort of thing to do deliberately on a quiet afternoon rather than as a side effect of moving a
|
||
service between systems.
|
||
|
||
**So there is a step before any of this: read the current environment out of the old system**, because
|
||
adoption means supplying those values and they live in its files today.
|
||
|
||
**And there is a failure worse than losing data, which is likelier.** A database image consumes its
|
||
password environment variable **only when its data directory is empty**. Everything here keeps its
|
||
data on a persistent directory, so the role holds whatever password it was created with, for ever.
|
||
Regenerate that variable and the application moves on while the database does not — permanently,
|
||
because nothing reconciles it. Eight modules are in that state today, working only because nobody
|
||
has regenerated their credential since their data directory was created.
|
||
|
||
*Where the detail lives:* this is operational and names machines, so it is in the mesh's own
|
||
knowledge base rather than here — `migration/where-service-data-lives`, which surveys where every
|
||
service's data actually sits and what each stop or removal would cost, and
|
||
`troubleshooting/db-password-frozen-at-first-init` for the lockout itself. **This document says the
|
||
rule; those say the specifics.**
|
||
|
||
*Corrected 2026-08-31 — an earlier version of this paragraph made that sound more dangerous than it
|
||
is.* A sealed secret is not unreadable; it is sealed **to the node**, which holds the private half
|
||
and writes the plaintext into the module's own file. The value is there, on the machine, as an
|
||
ordinary file. What does not exist is a way to ask *the mesh* what a secret is, and there is no
|
||
reveal command, because a mesh that can reveal a secret is a mesh that holds one.
|
||
|
||
## Where it starts, and what that costs
|
||
|
||
**On the node holding all the production data**, because that is where the services being
|
||
converted actually are.
|
||
|
||
Recorded plainly rather than argued with: this is the highest-risk order available. Everything
|
||
proven so far was proven on machines that could be destroyed and raised again, and the first real
|
||
exercise of the conversion will be on the one machine where a mistake is not recoverable. Nothing
|
||
about the lab work transfers automatically — a scenario proves the mechanism, not the state on
|
||
that machine.
|
||
|
||
**What makes it survivable is preparation rather than caution**: a restored backup before the
|
||
first step, one service at a time, and the previous arrangement left standing until the new one
|
||
has been read back. None of that is slower than the alternative, because the alternative includes
|
||
losing something.
|
||
|
||
## How the two systems hand over
|
||
|
||
*2026-08-31.* **The old system's brain is switched off; its services keep running.**
|
||
|
||
Not a migration and not a period of dual control. The old controller — provisioning, the
|
||
coordinator, the pipeline, the things that *decide* and *write* — is stopped. Every workload it
|
||
was managing goes on running exactly as it is, because nothing is managing it. Then the new mesh
|
||
takes ownership of them one at a time.
|
||
|
||
**Nothing is ever unassigned in the old system.** Unassigning is how it removes things, and
|
||
removing is how data is lost. The old system is never asked to take anything away; it is asked to
|
||
stop having opinions.
|
||
|
||
| | |
|
||
|---|---|
|
||
| **stopped, and disabled** | provisioning, the coordinator, environment and configuration sync, the pipeline — anything that decides or writes a file |
|
||
| **left alone entirely** | the units running the actual services: identity, mail, databases, the forge. They keep serving throughout |
|
||
| **never used** | unassign, remove, delete — any operation whose job is to take something away |
|
||
|
||
**Disabled, not merely stopped**, and this is the part that is easy to get wrong: those units are
|
||
enabled, so stopping them lasts until the machine reboots. A reboot mid-conversion would bring the
|
||
old controller back and it would resume regenerating managed files underneath the new one —
|
||
which is the one situation where two systems really would be fighting over the same machine.
|
||
|
||
**A service left running with nothing managing it is the safe state.** It has its data, its
|
||
configuration is already on disk, and nothing is going to change either. That is the whole trick:
|
||
the risk in a conversion is in the *managing*, not in the *running*.
|
||
|
||
**A brief interruption is acceptable. Losing data is not.** Where those two trade against each
|
||
other, the interruption wins every time — a service can be restarted, and there is no operation
|
||
that un-deletes a mail spool.
|
||
|
||
**The new host cannot remove what it did not put there.** Orphans are per-origin, so it only ever
|
||
removes resources it recorded itself. Services it has never been told about are not orphans to
|
||
it — they are simply not its business, which is what makes taking ownership one module at a time
|
||
safe.
|
||
|
||
## Phase 2 — the first real module
|
||
|
||
| # | task | done when |
|
||
|---|---|---|
|
||
| 2.1 | Port an **object store** module | it runs on the new mesh, serves a bucket to another module, and its credential rotates |
|
||
| 2.2 | Copy the data, and read it back through the service that owns it | the new location answers with what the old one holds |
|
||
| 2.3 | Point one dependent at it, old arrangement left standing | something real reads and writes through the new mesh's copy |
|
||
|
||
**Checkpoint, and it is a human one:** it runs for a week before anything else moves. The point of
|
||
going first is to find what Phase 1 missed, and a week is roughly how long that takes to show.
|
||
|
||
## Phase 3 — the modules that prove the shape
|
||
|
||
Each exercises something the first one does not.
|
||
|
||
| # | task | proves |
|
||
|---|---|---|
|
||
| 3.1 | An **identity provider** | a module that is itself a provider — the provides/requires chain, with consumers requiring it |
|
||
| 3.2 | A **forge** | a port claim against the machine's own daemon, and a module wanting both a database and an object store |
|
||
| 3.3 | A **mail system** | several containers as one module, a private network between them, and names that are not one-per-node |
|
||
|
||
**3.3 is the hardest thing in this document** and is deliberately last. If the declaration
|
||
language turns out to be insufficient, it says so here.
|
||
|
||
### Where Phase 3 actually stands — *2026-09-01*
|
||
|
||
All three have manifests. All three parse, resolve and plan. **None of them can start**, and the
|
||
two reasons are both filed rather than guessed at.
|
||
|
||
The **vocabulary held**. Nothing in 3.1–3.3 turned out to need a new shape: the identity provider,
|
||
the forge and the mail system are all expressible with what exists, including the mail system's
|
||
several containers on a private network — which was the one expected to break it. That is the
|
||
question this phase was designed to answer, and the answer is yes.
|
||
|
||
What did not hold was underneath the vocabulary:
|
||
|
||
- **[`022`](../../04-ISSUES/022-one-credential-per-node-per-provision-not-per-module/00-report.md)**
|
||
(fixed) — a credential belonged to a machine, so a node running several modules against one
|
||
database could not be planned. The refusal was loud on the provider and silent on the consumer.
|
||
- **[`023`](../../04-ISSUES/023-a-consumer-cannot-build-a-connection-string/00-report.md)** (open)
|
||
— a consumer gets its password and still cannot connect: the user name is invented by the
|
||
provisioner and recorded nowhere, and the bound values cannot reach a configuration file.
|
||
|
||
A third fault was in the manifests themselves rather than the design: each declared a secret at a
|
||
path named `.env` and read it as one, when a sealed file holds a password and nothing else. They
|
||
parsed and resolved and could never have worked, which is what a manifest checked only by a parser
|
||
buys. Two tests now refuse both halves of it.
|
||
|
||
### 3.1 needs a program, not a decision — *2026-09-01*
|
||
|
||
023 is fixed, and with it the two design faults are gone. What stands between 3.1 and a running
|
||
identity provider is now one concrete thing: **the realm provisioner does not exist.**
|
||
|
||
Its manifest named an image — `mesh-provision-keycloak` — that nothing builds and no program
|
||
backs. That has been removed rather than left standing, because a manifest describing a program
|
||
nobody wrote is the same mistake as the credential files that could never be read: it parses, it
|
||
resolves, and it could never work.
|
||
|
||
So Keycloak's manifest now says what is true today — a server the mesh runs, with its database
|
||
and its admin credential, both reaching it in a shape it can read. It no longer claims to provide
|
||
`oidc-client`, which means a consumer asking for one is **refused by name at plan time** rather
|
||
than resolving cleanly and waiting for a client nothing will create.
|
||
|
||
The provisioner is the same shape as the two that exist: it reads what the mesh granted and
|
||
reconciles a realm and a client per consumer. **It should be written against a real Keycloak in
|
||
the lab**, not from the API documentation — the object store's took three corrections that only a
|
||
running server produced.
|
||
|
||
The forge (3.2) and the mail system (3.3) need no provisioner and are not blocked on this.
|
||
|
||
### And they could not have run anyway — *2026-09-01*
|
||
|
||
Every one of the five named a container image that does not exist: sixty-four zeros where a digest
|
||
belongs, eighteen times over
|
||
([`025`](../../04-ISSUES/025-a-module-must-pin-a-digest-and-nothing-produces-one/00-report.md)).
|
||
They parsed, resolved and composed into a declaration a host accepts, and every one would have
|
||
stopped on the machine at the moment of fetching.
|
||
|
||
Nothing caught it because nothing could. A host checks the *shape* of a reference and no more —
|
||
verifying a digest exists means reaching a registry, which is the one thing a host must never have
|
||
to do. The refusal now sits where a declaration is composed instead, which is the last moment
|
||
before a machine sees one.
|
||
|
||
Twelve are pinned to real images. Two further faults surfaced only by pinning for real: the mail
|
||
system's seven images named repositories that **do not exist**, because it publishes to a
|
||
different registry than assumed, and one of the seven had been renamed upstream.
|
||
|
||
**The forge now runs**, on a database another module provides, with a password it did not choose
|
||
and a connection string it could not have written. That is the first of these descriptions to be
|
||
started rather than planned, and it exercises everything the credential work added.
|
||
|
||
What is still missing is the mechanism: nothing turns a tag into a digest as part of the mesh's
|
||
own work, so it was done by hand. Asking a registry takes about a second and pulls nothing, which
|
||
removes the main argument for leaving it undone.
|
||
|
||
## The conversion is done by hand, and that is a decision
|
||
|
||
*2026-08-31.* **Moving from the current system to this one is a person at a command line, working
|
||
through it.** Not a migration program, not a converter, not a period of dual-writing.
|
||
|
||
**What that removes from this plan is larger than what it adds.** Nothing below needs an importer,
|
||
a translation layer, a compatibility shim, or a mechanism for keeping two systems agreeing while
|
||
both are live — and every one of those is a thing somebody would otherwise reasonably build, use
|
||
once, and maintain for a year. The modules are the input; a person reads what one does today and
|
||
writes what it declares tomorrow.
|
||
|
||
**It also changes what "safe" means for the system being retired.** A fix to it has to be safe on
|
||
its own, because there is no careful rollout to sequence it into: the thing is being switched off
|
||
by hand, not managed into retirement. A change needing three steps in the right order is a change
|
||
that will be half-applied.
|
||
|
||
**And it is why the checkpoints below are weeks rather than gates.** Nothing enforces the order —
|
||
a person does — so the value of the sequence is entirely in what each step teaches before the next
|
||
one starts.
|
||
|
||
## Phase 4 — switch the old registry off
|
||
|
||
| # | task | done when |
|
||
|---|---|---|
|
||
| 4.1 | Move the remainder, by hand, a module at a time | nothing is assigned in the old system that is not assigned in the new one |
|
||
| 4.2 | The old one authoritative for nothing | a change to any module goes through the new mesh only |
|
||
| 4.3 | Switch it off | it is stopped, and nothing notices |
|
||
|
||
**4.3 is a day's work and the phases above it are not.** Naming it as a phase is what stops it
|
||
being mistaken for the goal.
|
||
|
||
## Sequencing
|
||
|
||
- **1 before 2.** Attempting a module without the vocabulary it needs produces a workaround, and a
|
||
workaround in a manifest is a design decision taken by whoever was in a hurry.
|
||
- **2 before 3, with the week.** Moving three modules before running one is how three modules
|
||
acquire the same defect.
|
||
- **3.3 last.** It is the only one that may send work back into the declaration language.
|
||
- **4 cannot start early, and there is no partial credit.** A registry still authoritative for one
|
||
module is still running.
|
||
|
||
## How this list is kept true
|
||
|
||
*This section exists because the document it replaces was wrong for weeks and nothing said so.*
|
||
|
||
**A claim here is counted, not reasoned.** The review of 2026-08-31 found a bundle described as
|
||
carrying two images that carries three, a bootstrap described as needing six shapes that uses
|
||
four, and ten documents calling themselves `designed` while naming lab-proven code. Each was
|
||
produced by describing the system from its design instead of reading it.
|
||
|
||
**A phase is done when the lab says so**, and the lab keeps a receipt of when it last ran and
|
||
against which commits. A phase marked done here whose assertions have not run is a claim about the
|
||
past.
|
||
|
||
**What is not proven gets said.** Phase 0 is done and its limitation is written into it. A list
|
||
that records only progress becomes a list nobody believes.
|
||
|
||
## Rules of engagement
|
||
|
||
Unchanged from the previous version: they were about how work is done rather than what the work
|
||
is.
|
||
|
||
### Autonomous by default
|
||
|
||
Read anything, measure anything, query read-only. Create branches, write code and tests, run the
|
||
suites, and write or update documents here.
|
||
|
||
### Always stop and ask
|
||
|
||
- **destroying or overwriting data** — dropping a table, deleting a provision, rotating a live
|
||
credential, removing a module from a node
|
||
- **merging anything** — every merge is a human checkpoint, without exception
|
||
- **anything touching a machine outside the lab**, including a configuration change that restarts
|
||
something people are using
|
||
- **a decision the records do not already answer** — record the question rather than picking and
|
||
moving on
|
||
- **any change to [`00-META`](../../00-META/)** — it is stable by nature
|
||
|
||
### Definition of done for every task
|
||
|
||
1. tests written **and failing first**, then passing
|
||
2. typecheck clean in every package the change touches
|
||
3. the behaviour demonstrated **in the lab, on real machines** — not asserted
|
||
4. documents here updated if the task changed or answered anything recorded
|
||
5. delivered, and the **effect** verified — not that a pipeline was green
|
||
|
||
### Non-negotiables
|
||
|
||
- **Never edit mesh-managed files on disk.** Use the thing that owns the file.
|
||
- **Never write to a production database directly.** Migrations for schema, application code for
|
||
data.
|
||
- **Every schema change ships twice** — consolidated schema *and* an incremental migration.
|
||
- **Expand, then contract.** Add the new shape, migrate, verify, and only then remove the old one.
|
||
- **A green pipeline proves transport, not effect.**
|
||
|
||
## What "done" looks like
|
||
|
||
The old registry is off. Every module runs on the new mesh, declared rather than scripted. A
|
||
machine that fails says what it could not do. And the number of modules grows when the work does,
|
||
not when the platform needs somewhere to put something.
|