Files
hq/01-RESEARCH/008-delivery-coordinator/00-overview.md
T
jschoubben 4ab8a0507f Delivery is reconciliation, not a pipeline; research 008 closes
Jochen: don't rebuild the current coordinator, use it as a pitfall list. That
reframed the last open question rather than answering it.

0058 stopped deploy being a stage that pushes to nodes, and said plainly what
it did not fix: detection. A merge that created no pipeline, and nothing said
so. That is not a defect in the detector -- it is what happens when correctness
depends on an event ARRIVING.

0063 applies 0058's move one level up. The control plane holds what source
exists and what has been built from it, and builds the difference. A change
becomes a build because source is ahead of artifacts, which is a comparison
answerable at any moment. An event makes it fast; nothing makes it necessary,
so a missed webhook costs latency and cannot cost correctness.

The mesh becomes one idea at two layers: the control plane reconciles artifacts
against source, the host reconciles machine state against declarations. The
pipeline as a state machine disappears, and with it the stage list that a
verify step was once omitted from.

That reframing answered the three questions still open in 008, so it graduates
with all six closed. A deployed state is two comparisons rather than an event.
A verdict is about an ARTIFACT and gates whether it may be declared -- sharper
than the question expected. And "before self-hosting" mostly dissolves, because
a reconciler needs source and artifacts as bindings where a pipeline's stages
name their targets.

Four costs recorded, and one is a real risk rather than a trade: a reconciler
that cannot reach its target retries forever, and without something noticing,
the failure is silence -- the exact fault this removes, reintroduced elsewhere.
Also named: the run identity people actually use is lost, and "did my change go
out?" needs a replacement or this will be worse to live with than what it
replaces, whatever its properties.
2026-08-28 02:59:42 +02:00

6.1 KiB

status, initiated, touches, became
status initiated touches became
graduated 2026-08-23
03-DESIGN/00-as-is/04-delivery.md
02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md
02-DECISIONS/0013-an-artifact-is-build-output.md
01-RESEARCH/006-mesh-from-scratch/code-skeleton.md
02-DECISIONS/0058-delivery-ends-in-a-declaration.md
02-DECISIONS/0063-delivery-is-reconciliation-not-a-pipeline.md

008 — The coordinator: a change checked in becomes a deployed state

What is being investigated

The mesh's own continuous delivery: a change is committed, and the mesh ends up in the state that change describes — across every node the change touches, with a verdict that says whether it worked.

The coordinator is what orchestrates that, and it is the mesh's most consequential machinery: everything reaches every node through it.

Why now

The as-is record (03-DESIGN/00-as-is/04-delivery.md) names problems that are structural rather than incidental:

  • A green pipeline proves transport, not effect. The stages report that a message was dispatched and accepted, which is not the same as the thing running, correct, or present. This is the mesh's single most consistent failure shape.
  • Detection is the most fragile input. A merge that creates no pipeline, with nothing saying so, is the characteristic bad outcome — and it has happened for reasons unrelated to the change.
  • The fan-out point is asymmetric. The build node has already passed two silos when work fans out, and code that knew only about the first parked it forever while every other node deployed cleanly.
  • There is no end-to-end coverage. The harness has not built since 2026-06-04 (04-ISSUES/005).

Research 006 adds a requirement the current design does not have: the coordinator must work before the mesh is self-hosting, when source and artifacts come from outside, and keep working across the transition to self-hosted providers.

What it became

Closed 2026-08-28. All six questions are answered, by two records, and the second exists because the first was honest about what it did not fix.

Does the coordinator dispatch stages, or converge nodes on a declaration? — Converge. ADR 0058: a pipeline ends when the declaration is updated, and the host applies it and reads back — so the reporter is the applier.

Does the three-silo split survive? — Yes, with the third redefined. The cardinality observation holds; the third silo is not a stage any more.

How does a change become a pipeline, reliably? — It does not become a pipeline at all. ADR 0063 applies 0058's move one level up: the control plane holds what source exists and what has been built, and builds the difference. An event makes it fast; nothing makes it necessary. The failures this effort catalogued — a truncated commit list, a broken path match — become latency rather than silence.

What is a deployed state? — Two comparisons, not an event. Does every node's reported state match what is declared, and is what is declared built from current source? A milestone can be claimed by something that did not check; a comparison cannot.

What produces a verdict, and what is it about? — An artifact, and it gates eligibility. The mesh must not converge onto something broken, so an artifact may be declared only once the lab has judged it fit. Sharper than the question expected: a verdict is a property an artifact has, not a report about a run.

How does delivery work before self-hosting? — It mostly stops being a question. A reconciler reads source and writes artifacts; where those live is a binding, external first and internal later. The transition looked hard because a pipeline's stages name their targets.

What this effort was right about

Its first question — what is a deployed state, and how does the mesh know it is in one — was marked "everything follows from this", and everything did. Both records above are answers to it: 0058 makes the applier the reporter, and 0063 makes currency a comparison. The effort put the load-bearing question first.

What is NOT closed by this

ADR 0063 names four costs and one of them is a real risk rather than a trade: a reconciler that cannot reach its target retries forever, and without something that notices, the failure is silence — which is the fault this effort exists to catalogue, reintroduced in a new place. That belongs to observability and it is not designed.

The questions (all answered above)

Question Why it matters
What is a deployed state, and how does the mesh know it is in one? Everything follows from this. If a stage reports transport, "deployed" is a claim nobody checked. A desired-state model with reconciliation gives a different answer from a job-completion model.
Does the coordinator dispatch stages, or converge nodes on a declaration? The current model is a state machine over stages. The alternative is that a node is told what should be true and reports what is. The second makes drift visible; the first cannot see it.
How does a change become a pipeline, reliably? Detection has failed for reasons unrelated to the change, silently.
What produces a verdict, and what is it a verdict about? Ties to the lab (ADR 0016) and to a module carrying its own assertions.
How does delivery work before self-hosting, and across the transition? From research 006: source and artifacts start external and are re-bound to internal providers. The coordinator has to be indifferent to which.
Does the three-silo split survive the artifact/part split? ADR 0014 is cardinality-driven, and research 006 renames the thing the cardinality is about.