Ordering needed no change for the third time running — resources apply in the order declared and nothing sorts them — and is now asserted, because sorting them for any sensible reason would have passed every other test. Separates ordering from readiness, which the task had run together: a container started is not a container ready. Nothing waits, and what needs something usable retries. That is deliberate and more robust than start ordering, since a dependency can restart long after apply. The network was the first thing in Phase 1 that genuinely needed building, and the first that needed a decision: 0029 records why a shape rather than an action, and the vocabulary is nine.
12 KiB
layer, status, code, updated, decisions
| layer | status | code | updated | decisions | |||||
|---|---|---|---|---|---|---|---|---|---|
| to-be | designed | 2026-08-31 |
|
Work breakdown — replacing what provisions the mesh
Rewritten 2026-08-31. The previous version planned a decomposition of the existing system in place: extract contexts, convert modules to declared features, shrink its shared library. That is not what is being done — a replacement is being built beside it, and the old plan's Phase 0 was the only part that survived contact with it. So the document that was supposed to say what happens next had been describing work on a system being retired.
The goal, in one sentence
Modules move to the new mesh one at a time, until the old registry can be switched off.
Everything below is ordered by what that requires. Nothing here is a rewrite of the old system; its modules are the input.
Phase 0 — a mesh that runs — done
Not the code exists. Twenty-two assertions on real machines in the lab, each confirmed to fail when the behaviour is removed (ADR 0016, ADR 0017).
| what is proven | |
|---|---|
| a mesh comes into being | a bare machine becomes one; others join with nothing but a token |
| credentials | delivered to both ends with the mesh holding neither; rotated so the old one stops working |
| declarations survive reality | a stopped machine is waited for; one that fell behind catches up unnamed; unassigning takes away exactly what it should; what the mesh says nothing about is left alone |
| failure is legible | a machine that cannot do what it was told is named, with why |
| the mesh runs itself | its own artifact store, and a builder that is a module the mesh assigns |
| names and reachability | internal names, wildcards under a machine, containers reaching other machines, certificates the mesh issued, filtering that matches exactly what was declared |
| delivery | a new commit reaches a machine already running the old one |
| model access | answered by a record, with a key the mesh cannot read |
What Phase 0 does not prove, and it is the important sentence in this document: every module
exercised above was written to test the mechanism. No module from the existing system has ever
run on this. The vocabulary was shaped by the things used to test it — the same fault as a
fixture agreeing with the code it checks
(04-ISSUES/005), at the
scale of a design.
Phase 1 — the vocabulary a real module needs
Found by taking real modules and asking what they would require. Each is a gap in what can be expressed, not a defect in what is built.
| # | task | done when |
|---|---|---|
| seven assertions against a real store | ||
| two sessions on one machine, different licences, each its own key | ||
| the shape is created and removed; ordering was already there, and is now asserted | ||
| 1.4 | Public certificate issuance | a name reachable from outside is served with a certificate from a public authority, obtained against a staging endpoint unless told otherwise (04-ISSUES/004) |
1.3 and 1.4 block later ones and are listed now so they are not met as surprises. 1.3 is what a mail system needs and nothing else so far does.
Checkpoint: each is demonstrated in the lab before the module needing it is attempted.
1.2, and the same surprise twice
A binding is per module per machine, and the two sessions are two modules — the same mechanism
in different context roots, and a context root is what a module delivers. So (node, module)
already names them apart, and nothing needed adding.
14-model-access.md had called per-module-per-machine a step toward it and
not it, which is true of a worker — many run on one machine from one module — and not true of
a session, of which there is one per node and one for the mesh.
1.3, and the first one that needed building
Ordering was already there — the apply loop sorts nothing, so a module says this before that by writing it first. Untested until now, and the kind of property a later change breaks silently. Worth separating from readiness: a container started is not a container ready, and nothing waits. What needs something usable retries, which is what both provisioners do and is the better answer anyway, because a dependency can restart long after everything was applied.
The network was a real gap, and the first thing in Phase 1 that needed a decision. Adding a
shape widens what a compromised control plane can express, so
ADR 0029
records why this one is worth it: an action could create a network and nothing could ever
remove it, because an action leaves no footprint the host can undo. The vocabulary is nine.
Three tasks in a row that were already possible. Both were written from the design rather than from the code, which is the review's finding arriving in the plan: a claim here is counted, not reasoned. The remaining Phase 1 items should be checked against the code before being started, not after.
1.1, and what it turned out to be
Done 2026-08-31. Worth recording because the task was not the one written down.
The control plane special-cases nothing. provides, requires, contributes and grants
are name-agnostic — asking for a bucket needed no change to the mesh at all. What was missing was
a provider, and the last step where something on the machine turns a delivered secret into a key
that works. So "add an object-store provision" was never mesh work.
The provision is s3-bucket: a consumer's code is written against the S3 API and swapping one
store for another does not break it, so by
ADR 0027 the name
says the protocol. A database is the other case, and names the engine.
One assertion here that a database does not need. One PostgreSQL server holds separate databases and the product enforces the boundary; one object store holds every bucket behind one endpoint, so a consumer cannot reach another consumer's bucket is a policy somebody wrote — and a policy granting everything would pass every other test. What is asserted is what the policy does not say.
Phase 2 — the first real module
| # | task | done when |
|---|---|---|
| 2.1 | Port an object store module | it runs on the new mesh, serves a bucket to another module, and its credential rotates |
| 2.2 | Run it beside the existing one | both exist; nothing depends on the new one yet |
| 2.3 | Move one dependent onto it | something real reads and writes through the new mesh's copy |
Checkpoint, and it is a human one: it runs for a week before anything else moves. The point of going first is to find what Phase 1 missed, and a week is roughly how long that takes to show.
Phase 3 — the modules that prove the shape
Each exercises something the first one does not.
| # | task | proves |
|---|---|---|
| 3.1 | An identity provider | a module that is itself a provider — the provides/requires chain, with consumers requiring it |
| 3.2 | A forge | a port claim against the machine's own daemon, and a module wanting both a database and an object store |
| 3.3 | A mail system | several containers as one module, a private network between them, and names that are not one-per-node |
3.3 is the hardest thing in this document and is deliberately last. If the declaration language turns out to be insufficient, it says so here.
Phase 4 — switch the old registry off
| # | task | done when |
|---|---|---|
| 4.1 | Move the remainder | nothing is assigned in the old system that is not assigned in the new one |
| 4.2 | Run in parallel, the old one authoritative for nothing | a change to any module goes through the new mesh only |
| 4.3 | Switch it off | it is stopped, and nothing notices |
4.3 is a day's work and the phases above it are not. Naming it as a phase is what stops it being mistaken for the goal.
Sequencing
- 1 before 2. Attempting a module without the vocabulary it needs produces a workaround, and a workaround in a manifest is a design decision taken by whoever was in a hurry.
- 2 before 3, with the week. Moving three modules before running one is how three modules acquire the same defect.
- 3.3 last. It is the only one that may send work back into the declaration language.
- 4 cannot start early, and there is no partial credit. A registry still authoritative for one module is still running.
How this list is kept true
This section exists because the document it replaces was wrong for weeks and nothing said so.
A claim here is counted, not reasoned. The review of 2026-08-31 found a bundle described as
carrying two images that carries three, a bootstrap described as needing six shapes that uses
four, and ten documents calling themselves designed while naming lab-proven code. Each was
produced by describing the system from its design instead of reading it.
A phase is done when the lab says so, and the lab keeps a receipt of when it last ran and against which commits. A phase marked done here whose assertions have not run is a claim about the past.
What is not proven gets said. Phase 0 is done and its limitation is written into it. A list that records only progress becomes a list nobody believes.
Rules of engagement
Unchanged from the previous version: they were about how work is done rather than what the work is.
Autonomous by default
Read anything, measure anything, query read-only. Create branches, write code and tests, run the suites, and write or update documents here.
Always stop and ask
- destroying or overwriting data — dropping a table, deleting a provision, rotating a live credential, removing a module from a node
- merging anything — every merge is a human checkpoint, without exception
- anything touching a machine outside the lab, including a configuration change that restarts something people are using
- a decision the records do not already answer — record the question rather than picking and moving on
- any change to
00-META— it is stable by nature
Definition of done for every task
- tests written and failing first, then passing
- typecheck clean in every package the change touches
- the behaviour demonstrated in the lab, on real machines — not asserted
- documents here updated if the task changed or answered anything recorded
- delivered, and the effect verified — not that a pipeline was green
Non-negotiables
- Never edit mesh-managed files on disk. Use the thing that owns the file.
- Never write to a production database directly. Migrations for schema, application code for data.
- Every schema change ships twice — consolidated schema and an incremental migration.
- Expand, then contract. Add the new shape, migrate, verify, and only then remove the old one.
- A green pipeline proves transport, not effect.
What "done" looks like
The old registry is off. Every module runs on the new mesh, declared rather than scripted. A machine that fails says what it could not do. And the number of modules grows when the work does, not when the platform needs somewhere to put something.