Three things from walking a real dev cycle through 0063, all of which Jochen caught by pushing on where I had glossed. 0064 -- a build edge is a third kind. Research 011 established presence and instantiation, and both are RUNTIME edges: they answer what a module needs in order to run. Delivery needs a different question -- what has to be rebuilt when this changes -- and that relationship is fixed inside an artifact rather than negotiated when it runs. So the graph as designed could not drive delivery, which is the real reason 0063 was not approvable. It is derived rather than declared, read from what a module actually imports, because a declared list and the imports it describes drift and the imports are the true ones. The runtime edges stay declared, and that asymmetry is not an inconsistency: a runtime edge is an intention somebody has, a build edge is a fact about code that exists. It also makes design quality measurable. A module with many inbound build edges is one whose every change is expensive, and the current shared library is exactly that -- nobody could see it because nothing drew the edges. 0065 -- the core library is the mesh's domain. Jochen disagreed with 0030's "types, not behaviour" and was right: that guard is aimed at the wrong thing. A library everything depends on is a hub whether it holds types or code, and the fan-in is what makes a change expensive. So types ship with the module that owns them -- trading one wide edge for several narrow ones -- and the core library holds what is true of the mesh regardless of context, which research 011 already found: a module, a node, an assignment. The test is "would this still mean the same thing in a context that had never heard of the one it came from". A node does; a pipeline stage does not. Domain-driven is the point rather than the label: "who else might want this" always answers yes, which is how the current one grew. And it changes the check for the better. "The build output contains no runtime code" would have enforced a rule now withdrawn. Inbound build edges is a measurement rather than a prohibition, and it is visible while a hub is forming rather than after. 0063 revised on both counts, plus a third: I had written "the lab judges it" as though that were a step. A lab run takes tens of seconds, occupies a VM, and fails for environmental reasons -- and a shared-library change produces dozens. One expensive non-deterministic gate fails both ways, and neither failure looks like itself. Verdicts are now tiered, and a run that failed environmentally is explicitly not a verdict. 0063 also now carries what must exist before it can be implemented, rather than leaving that to be discovered.
7.1 KiB
status, date, deciders, reconstructed, extends
| status | date | deciders | reconstructed | extends |
|---|---|---|---|---|
| proposed | 2026-08-28 | jochen | false | 0058-delivery-ends-in-a-declaration.md |
63. Delivery is reconciliation, not a pipeline
Context
ADR 0058 stopped deploy being a stage that pushes to nodes: the control plane says what should be true, the host reconciles, and the thing that reports is the thing that did the work. It fixed the third silo and left the first two as they were — and it said plainly what it did not fix:
Detection stays the fragile input, and this does not fix it. A merge that created no pipeline, and nothing said so is upstream of everything here and is untouched.
That is not a defect in the detector. It is what happens when a system's correctness depends on an event arriving. The as-is records the ways it has failed — a webhook truncating its commit list on a large merge, a forge address whose port broke module-path matching — and the shape is always the same: nothing happened, and nothing said so.
The same move that fixed deploy fixes this, applied one level up. This record is not a better detector. It is the removal of detection as a load-bearing mechanism.
And this is not the current coordinator repaired. The existing pipeline is a state machine over stages; what follows is not that with better inputs. The old system's value here is as a catalogue of the ways this can fail, and it has been used for exactly that.
Decision
The control plane holds what source exists and what has been built from it, and builds the difference.
A change becomes a build because source is ahead of artifacts — a comparison, answerable at any moment — rather than because a message arrived.
An event makes it fast. Nothing makes it necessary. A push notification is an optimisation that lowers latency; a missed one costs latency and cannot cost correctness. That is the same property the host's drift timer has, and it is the whole point of both.
So the mesh is one idea at two layers:
| reconciles | against | |
|---|---|---|
| the control plane | artifacts | source |
| the host | machine state | declarations |
What disappears: the pipeline as a state machine. There is no stage list something can be omitted from — which is how a verify stage was built and never scheduled — and no run to lose.
What this answers
Research 008 asked six questions. Two were answered by ADR 0058; this answers the rest.
What is a deployed state? Not an event — two comparisons, both answerable on demand: does every node's reported state match what is declared, and is what is declared built from current source? A milestone can be claimed by something that did not check. A comparison cannot.
What produces a verdict, and what is it about? An artifact, and it gates eligibility. The mesh must not converge onto something broken, so an artifact may be declared only once something has judged it fit. The lab is what judges. This is sharper than the question expected: a verdict is not a report about a run, it is a property an artifact does or does not have.
How does delivery work before self-hosting? It mostly stops being a question. A reconciler needs to read source and write artifacts; where those live is a binding, external at first and internal later. A pipeline has stages that name their targets, which is why the transition looked hard.
What survives from ADR 0014
ADR 0014 is a decision about cardinality, and the cardinality observation is right and unchanged: building is per module, publishing is per module, and what happens on nodes is per node. What changes is that those are no longer three silos of a job. They are steps of reconciling one artifact, and the third is not a step at all any more (ADR 0058).
What this costs
Named because each is a way this can go wrong, and a decision that lists none has not been examined.
- "Is this artifact current?" must be answerable without building it. A commit recorded against each artifact does it, and that record becomes load-bearing: wrong, and the mesh either rebuilds forever or never rebuilds at all.
- Rebuild storms are real and mostly behaviourally empty. One shared-library commit invalidates nearly everything, and most of those rebuilds produce artifacts that do the same thing they did before — so the fleet is redeployed for no change in behaviour. Reproducible builds would stop the cascade at the first module whose output did not move; without them, the storm is in the declarations rather than the builds (ADR 0064).
- The run identity people actually use is lost. Did my change go out? is answerable today by opening a pipeline. With convergence there is no run to open, and something has to replace that — a query over the two comparisons above — or this will be worse to live with than what it replaces, whatever its properties.
- A loop that will not converge is harder to debug than a job that failed. A failed job stops and names its step. A reconciler that cannot reach its target retries forever, and without something that notices this has been trying for an hour, the failure is silence — which is the fault this record is removing, reintroduced in a new place. This is the real risk and it is not solved here.
What must exist first
Stated as a list because this record cannot be implemented without them, and saying so is better than discovering it:
- The module graph, including build edges (ADR 0064). Designed, not built. Without it there is no rebuild set and no ordering.
- A recorded input closure per artifact — its commit and the identity of everything it was built against — so is this current? is answerable without building.
- Something that notices a reconciler is not converging. Below, and the one that is a risk rather than a cost.
Consequences
- Detection stops being correctness and becomes latency. The specific faults the as-is records — a truncated commit list, a broken path match — become slow rather than silent.
- The coordinator is not ported. What replaces it is a comparison and a build, and the existing implementation informs it only as a list of things that went wrong.
- Two things must be cheap that are not yet designed: reading what source exists, and reading what has been built. Both are queries against providers the mesh will host, and both are on the path of everything above.
- Research 008 can close, which it could not before this.
References
- ADR 0058 — the same move, one level down.
- ADR 0014 — the cardinality that survives.
- ADR 0035 — why a comparison beats a claim.
00-as-is/04— the pitfalls this is designed against.