All three modules have manifests, all three parse, resolve and plan, and none of them can start. Worth writing down before it reads as progress or as failure, because it is neither. The vocabulary held. Nothing in 3.1–3.3 needed a new shape — including the mail system's several containers on a private network, which was the one expected to break it. That was the question this phase was designed to answer. What did not hold was underneath: 022, now fixed, and 023, open. Both are about credentials rather than about what a module can say. The third fault was in the manifests, not the design: a secret declared at a path named .env and read as one, when a sealed file holds a password and nothing else. That is what a manifest checked only by a parser buys, and it is why there are now two tests reading the manifests on disk. 023 is the whole of what remains before the identity provider runs.
22 KiB
layer, status, code, updated, decisions
| layer | status | code | updated | decisions | |||||
|---|---|---|---|---|---|---|---|---|---|
| to-be | designed | 2026-09-01 |
|
Work breakdown — replacing what provisions the mesh
Rewritten 2026-08-31. The previous version planned a decomposition of the existing system in place: extract contexts, convert modules to declared features, shrink its shared library. That is not what is being done — a replacement is being built beside it, and the old plan's Phase 0 was the only part that survived contact with it. So the document that was supposed to say what happens next had been describing work on a system being retired.
The goal, in one sentence
Modules move to the new mesh one at a time, until the old registry can be switched off.
Everything below is ordered by what that requires. Nothing here is a rewrite of the old system; its modules are the input.
Phase 0 — a mesh that runs — done
Not the code exists. Twenty-two assertions on real machines in the lab, each confirmed to fail when the behaviour is removed (ADR 0016, ADR 0017).
| what is proven | |
|---|---|
| a mesh comes into being | a bare machine becomes one; others join with nothing but a token |
| credentials | delivered to both ends with the mesh holding neither; rotated so the old one stops working |
| declarations survive reality | a stopped machine is waited for; one that fell behind catches up unnamed; unassigning takes away exactly what it should; what the mesh says nothing about is left alone |
| failure is legible | a machine that cannot do what it was told is named, with why |
| the mesh runs itself | its own artifact store, and a builder that is a module the mesh assigns |
| names and reachability | internal names, wildcards under a machine, containers reaching other machines, certificates the mesh issued, filtering that matches exactly what was declared |
| delivery | a new commit reaches a machine already running the old one |
| model access | answered by a record, with a key the mesh cannot read |
What Phase 0 does not prove, and it is the important sentence in this document: every module
exercised above was written to test the mechanism. No module from the existing system has ever
run on this. The vocabulary was shaped by the things used to test it — the same fault as a
fixture agreeing with the code it checks
(04-ISSUES/005), at the
scale of a design.
Phase 1 — the vocabulary a real module needs
Found by taking real modules and asking what they would require. Each is a gap in what can be expressed, not a defect in what is built.
| # | task | done when |
|---|---|---|
| seven assertions against a real store | ||
| two sessions on one machine, different licences, each its own key | ||
| the shape is created and removed; ordering was already there, and is now asserted | ||
| 1.4 | Public certificate issuance — built; one gap | ordering, the challenge and issuance are proven against a real authority; collecting the issued certificate is not (04-ISSUES/020) |
1.3 and 1.4 block later ones and are listed now so they are not met as surprises. 1.3 is what a mail system needs and nothing else so far does.
Phase 1 is closed with 1.4 partly open, deliberately. Two of its four items needed no code at all; the network shape was built; and certificates are configured correctly, order correctly, and are issued correctly — the client does not collect what the authority issued, against a server that exists to be a test server. That is filed rather than chased, because the remainder may say nothing about a real authority and the next thing to learn comes from moving a module rather than from a fourth lab run.
Checkpoint: each is demonstrated in the lab before the module needing it is attempted.
1.2, and the same surprise twice
A binding is per module per machine, and the two sessions are two modules — the same mechanism
in different context roots, and a context root is what a module delivers. So (node, module)
already names them apart, and nothing needed adding.
14-model-access.md had called per-module-per-machine a step toward it and
not it, which is true of a worker — many run on one machine from one module — and not true of
a session, of which there is one per node and one for the mesh.
1.3, and the first one that needed building
Ordering was already there — the apply loop sorts nothing, so a module says this before that by writing it first. Untested until now, and the kind of property a later change breaks silently. Worth separating from readiness: a container started is not a container ready, and nothing waits. What needs something usable retries, which is what both provisioners do and is the better answer anyway, because a dependency can restart long after everything was applied.
The network was a real gap, and the first thing in Phase 1 that needed a decision. Adding a
shape widens what a compromised control plane can express, so
ADR 0029
records why this one is worth it: an action could create a network and nothing could ever
remove it, because an action leaves no footprint the host can undo. The vocabulary is nine.
Three tasks in a row that were already possible. Both were written from the design rather than from the code, which is the review's finding arriving in the plan: a claim here is counted, not reasoned. The remaining Phase 1 items should be checked against the code before being started, not after.
1.1, and what it turned out to be
Done 2026-08-31. Worth recording because the task was not the one written down.
The control plane special-cases nothing. provides, requires, contributes and grants
are name-agnostic — asking for a bucket needed no change to the mesh at all. What was missing was
a provider, and the last step where something on the machine turns a delivered secret into a key
that works. So "add an object-store provision" was never mesh work.
The provision is s3-bucket: a consumer's code is written against the S3 API and swapping one
store for another does not break it, so by
ADR 0027 the name
says the protocol. A database is the other case, and names the engine.
One assertion here that a database does not need. One PostgreSQL server holds separate databases and the product enforces the boundary; one object store holds every bucket behind one endpoint, so a consumer cannot reach another consumer's bucket is a policy somebody wrote — and a policy granting everything would pass every other test. What is asserted is what the policy does not say.
Data is the constraint, and it outranks the order below
2026-08-31. The modules being converted run live services — identity, mail — and the data must survive every step. A data folder may move; it may never be lost.
One thing was found by asking this and is fixed (ADR 0030): the host deleted a directory and everything under it when the directory stopped being declared, which happens when a module is unassigned or a manifest is edited to move a data folder — the exact operation this plan needs. A directory holding anything the mesh did not put there is now kept and reported.
That is not a backup and must not be read as one. It stops the mesh destroying data. It does nothing about a disk, a mistaken command, or a service corrupting its own store.
So the rule for every step below: the data is copied, the copy is verified by reading it back through the service that owns it, and only then does anything point at the new location. Never moved and then checked. A backup nobody has restored is a belief, not a copy.
A module is adopted with the credentials it already has
2026-08-31. Nothing is rotated during the conversion. A service being adopted keeps the password it is already using, because minting a new one is how a running service stops being able to reach its own database in the middle of a migration.
The mesh has both paths and this needs the second:
| generate | a new secret, sealed to both ends. What a new module gets |
| accept | a value supplied from outside, sealed, plaintext discarded. What an adopted module gets |
Rotation is a separate act, afterwards, once everything works. The machinery for it is built and proven — a credential moving at both ends with the old one ceasing to work — and it is exactly the sort of thing to do deliberately on a quiet afternoon rather than as a side effect of moving a service between systems.
So there is a step before any of this: read the current environment out of the old system, because adoption means supplying those values and they live in its files today.
And there is a failure worse than losing data, which is likelier. A database image consumes its password environment variable only when its data directory is empty. Everything here keeps its data on a persistent directory, so the role holds whatever password it was created with, for ever. Regenerate that variable and the application moves on while the database does not — permanently, because nothing reconciles it. Eight modules are in that state today, working only because nobody has regenerated their credential since their data directory was created.
Where the detail lives: this is operational and names machines, so it is in the mesh's own
knowledge base rather than here — migration/where-service-data-lives, which surveys where every
service's data actually sits and what each stop or removal would cost, and
troubleshooting/db-password-frozen-at-first-init for the lockout itself. This document says the
rule; those say the specifics.
Corrected 2026-08-31 — an earlier version of this paragraph made that sound more dangerous than it is. A sealed secret is not unreadable; it is sealed to the node, which holds the private half and writes the plaintext into the module's own file. The value is there, on the machine, as an ordinary file. What does not exist is a way to ask the mesh what a secret is, and there is no reveal command, because a mesh that can reveal a secret is a mesh that holds one.
Where it starts, and what that costs
On the node holding all the production data, because that is where the services being converted actually are.
Recorded plainly rather than argued with: this is the highest-risk order available. Everything proven so far was proven on machines that could be destroyed and raised again, and the first real exercise of the conversion will be on the one machine where a mistake is not recoverable. Nothing about the lab work transfers automatically — a scenario proves the mechanism, not the state on that machine.
What makes it survivable is preparation rather than caution: a restored backup before the first step, one service at a time, and the previous arrangement left standing until the new one has been read back. None of that is slower than the alternative, because the alternative includes losing something.
How the two systems hand over
2026-08-31. The old system's brain is switched off; its services keep running.
Not a migration and not a period of dual control. The old control plane — provisioning, the coordinator, the pipeline, the things that decide and write — is stopped. Every workload it was managing goes on running exactly as it is, because nothing is managing it. Then the new mesh takes ownership of them one at a time.
Nothing is ever unassigned in the old system. Unassigning is how it removes things, and removing is how data is lost. The old system is never asked to take anything away; it is asked to stop having opinions.
| stopped, and disabled | provisioning, the coordinator, environment and configuration sync, the pipeline — anything that decides or writes a file |
| left alone entirely | the units running the actual services: identity, mail, databases, the forge. They keep serving throughout |
| never used | unassign, remove, delete — any operation whose job is to take something away |
Disabled, not merely stopped, and this is the part that is easy to get wrong: those units are enabled, so stopping them lasts until the machine reboots. A reboot mid-conversion would bring the old control plane back and it would resume regenerating managed files underneath the new one — which is the one situation where two systems really would be fighting over the same machine.
A service left running with nothing managing it is the safe state. It has its data, its configuration is already on disk, and nothing is going to change either. That is the whole trick: the risk in a conversion is in the managing, not in the running.
A brief interruption is acceptable. Losing data is not. Where those two trade against each other, the interruption wins every time — a service can be restarted, and there is no operation that un-deletes a mail spool.
The new host cannot remove what it did not put there. Orphans are per-origin, so it only ever removes resources it recorded itself. Services it has never been told about are not orphans to it — they are simply not its business, which is what makes taking ownership one module at a time safe.
Phase 2 — the first real module
| # | task | done when |
|---|---|---|
| 2.1 | Port an object store module | it runs on the new mesh, serves a bucket to another module, and its credential rotates |
| 2.2 | Copy the data, and read it back through the service that owns it | the new location answers with what the old one holds |
| 2.3 | Point one dependent at it, old arrangement left standing | something real reads and writes through the new mesh's copy |
Checkpoint, and it is a human one: it runs for a week before anything else moves. The point of going first is to find what Phase 1 missed, and a week is roughly how long that takes to show.
Phase 3 — the modules that prove the shape
Each exercises something the first one does not.
| # | task | proves |
|---|---|---|
| 3.1 | An identity provider | a module that is itself a provider — the provides/requires chain, with consumers requiring it |
| 3.2 | A forge | a port claim against the machine's own daemon, and a module wanting both a database and an object store |
| 3.3 | A mail system | several containers as one module, a private network between them, and names that are not one-per-node |
3.3 is the hardest thing in this document and is deliberately last. If the declaration language turns out to be insufficient, it says so here.
Where Phase 3 actually stands — 2026-09-01
All three have manifests. All three parse, resolve and plan. None of them can start, and the two reasons are both filed rather than guessed at.
The vocabulary held. Nothing in 3.1–3.3 turned out to need a new shape: the identity provider, the forge and the mail system are all expressible with what exists, including the mail system's several containers on a private network — which was the one expected to break it. That is the question this phase was designed to answer, and the answer is yes.
What did not hold was underneath the vocabulary:
022(fixed) — a credential belonged to a machine, so a node running several modules against one database could not be planned. The refusal was loud on the provider and silent on the consumer.023(open) — a consumer gets its password and still cannot connect: the user name is invented by the provisioner and recorded nowhere, and the bound values cannot reach a configuration file.
A third fault was in the manifests themselves rather than the design: each declared a secret at a
path named .env and read it as one, when a sealed file holds a password and nothing else. They
parsed and resolved and could never have worked, which is what a manifest checked only by a parser
buys. Two tests now refuse both halves of it.
023 is the whole of what remains before 3.1 runs.
The conversion is done by hand, and that is a decision
2026-08-31. Moving from the current system to this one is a person at a command line, working through it. Not a migration program, not a converter, not a period of dual-writing.
What that removes from this plan is larger than what it adds. Nothing below needs an importer, a translation layer, a compatibility shim, or a mechanism for keeping two systems agreeing while both are live — and every one of those is a thing somebody would otherwise reasonably build, use once, and maintain for a year. The modules are the input; a person reads what one does today and writes what it declares tomorrow.
It also changes what "safe" means for the system being retired. A fix to it has to be safe on its own, because there is no careful rollout to sequence it into: the thing is being switched off by hand, not managed into retirement. A change needing three steps in the right order is a change that will be half-applied.
And it is why the checkpoints below are weeks rather than gates. Nothing enforces the order — a person does — so the value of the sequence is entirely in what each step teaches before the next one starts.
Phase 4 — switch the old registry off
| # | task | done when |
|---|---|---|
| 4.1 | Move the remainder, by hand, a module at a time | nothing is assigned in the old system that is not assigned in the new one |
| 4.2 | The old one authoritative for nothing | a change to any module goes through the new mesh only |
| 4.3 | Switch it off | it is stopped, and nothing notices |
4.3 is a day's work and the phases above it are not. Naming it as a phase is what stops it being mistaken for the goal.
Sequencing
- 1 before 2. Attempting a module without the vocabulary it needs produces a workaround, and a workaround in a manifest is a design decision taken by whoever was in a hurry.
- 2 before 3, with the week. Moving three modules before running one is how three modules acquire the same defect.
- 3.3 last. It is the only one that may send work back into the declaration language.
- 4 cannot start early, and there is no partial credit. A registry still authoritative for one module is still running.
How this list is kept true
This section exists because the document it replaces was wrong for weeks and nothing said so.
A claim here is counted, not reasoned. The review of 2026-08-31 found a bundle described as
carrying two images that carries three, a bootstrap described as needing six shapes that uses
four, and ten documents calling themselves designed while naming lab-proven code. Each was
produced by describing the system from its design instead of reading it.
A phase is done when the lab says so, and the lab keeps a receipt of when it last ran and against which commits. A phase marked done here whose assertions have not run is a claim about the past.
What is not proven gets said. Phase 0 is done and its limitation is written into it. A list that records only progress becomes a list nobody believes.
Rules of engagement
Unchanged from the previous version: they were about how work is done rather than what the work is.
Autonomous by default
Read anything, measure anything, query read-only. Create branches, write code and tests, run the suites, and write or update documents here.
Always stop and ask
- destroying or overwriting data — dropping a table, deleting a provision, rotating a live credential, removing a module from a node
- merging anything — every merge is a human checkpoint, without exception
- anything touching a machine outside the lab, including a configuration change that restarts something people are using
- a decision the records do not already answer — record the question rather than picking and moving on
- any change to
00-META— it is stable by nature
Definition of done for every task
- tests written and failing first, then passing
- typecheck clean in every package the change touches
- the behaviour demonstrated in the lab, on real machines — not asserted
- documents here updated if the task changed or answered anything recorded
- delivered, and the effect verified — not that a pipeline was green
Non-negotiables
- Never edit mesh-managed files on disk. Use the thing that owns the file.
- Never write to a production database directly. Migrations for schema, application code for data.
- Every schema change ships twice — consolidated schema and an incremental migration.
- Expand, then contract. Add the new shape, migrate, verify, and only then remove the old one.
- A green pipeline proves transport, not effect.
What "done" looks like
The old registry is off. Every module runs on the new mesh, declared rather than scripted. A machine that fails says what it could not do. And the number of modules grows when the work does, not when the platform needs somewhere to put something.