Files
hq/01-RESEARCH/008-delivery-coordinator/00-overview.md
T
jschoubben 77f3a4cea7 Consolidate: 65 decision records to 23
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.

  the node host          8 -> 1    applies not decides, depends on nothing,
                                   per operating system, root service, the
                                   launcher, episodic, what a declaration is,
                                   actions from the bundle only
  a node and how it joins 4 -> 1   what a node is, joining, the link as
                                   security boundary, the enrolment token
  modules and the graph   7 -> 1   everything is a module, no domain modules,
                                   three edges, provisioning, the core library
  substrate and control   6 -> 1   the test, seven contexts, one control plane,
    plane                          the authority is not a database, the named
                                   products, the pinned bundle
  connectivity            3 -> 1   a route is a grant, reachability declared,
                                   filter rules
  delivery                5 -> 1   reconciliation not a pipeline, artifacts,
                                   the three silos, a failed step, the verdict
  the lab                 5 -> 1   (earlier)
  how this repository     10 -> 1  (earlier)
    works

Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.

The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.

The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
2026-08-28 20:03:24 +02:00

5.8 KiB

status, initiated, touches, became
status initiated touches became
graduated 2026-08-23
03-DESIGN/00-as-is/04-delivery.md
02-DECISIONS/0058-delivery.md
02-DECISIONS/0058-delivery.md
01-RESEARCH/006-mesh-from-scratch/code-skeleton.md
02-DECISIONS/0058-delivery.md
02-DECISIONS/0058-delivery.md

008 — The coordinator: a change checked in becomes a deployed state

What is being investigated

The mesh's own continuous delivery: a change is committed, and the mesh ends up in the state that change describes — across every node the change touches, with a verdict that says whether it worked.

The coordinator is what orchestrates that, and it is the mesh's most consequential machinery: everything reaches every node through it.

Why now

The as-is record (03-DESIGN/00-as-is/04-delivery.md) names problems that are structural rather than incidental:

  • A green pipeline proves transport, not effect. The stages report that a message was dispatched and accepted, which is not the same as the thing running, correct, or present. This is the mesh's single most consistent failure shape.
  • Detection is the most fragile input. A merge that creates no pipeline, with nothing saying so, is the characteristic bad outcome — and it has happened for reasons unrelated to the change.
  • The fan-out point is asymmetric. The build node has already passed two silos when work fans out, and code that knew only about the first parked it forever while every other node deployed cleanly.
  • There is no end-to-end coverage. The harness has not built since 2026-06-04 (04-ISSUES/005).

Research 006 adds a requirement the current design does not have: the coordinator must work before the mesh is self-hosting, when source and artifacts come from outside, and keep working across the transition to self-hosted providers.

What it became

Closed 2026-08-28. All six questions are answered, by two records, and the second exists because the first was honest about what it did not fix.

Does the coordinator dispatch stages, or converge nodes on a declaration? — Converge. ADR 0058: a pipeline ends when the declaration is updated, and the host applies it and reads back — so the reporter is the applier.

Does the three-silo split survive? — Yes, with the third redefined. The cardinality observation holds; the third silo is not a stage any more.

How does a change become a pipeline, reliably? — It does not become a pipeline at all. ADR 0058 applies 0058's move one level up: the control plane holds what source exists and what has been built, and builds the difference. An event makes it fast; nothing makes it necessary. The failures this effort catalogued — a truncated commit list, a broken path match — become latency rather than silence.

What is a deployed state? — Two comparisons, not an event. Does every node's reported state match what is declared, and is what is declared built from current source? A milestone can be claimed by something that did not check; a comparison cannot.

What produces a verdict, and what is it about? — An artifact, and it gates eligibility. The mesh must not converge onto something broken, so an artifact may be declared only once the lab has judged it fit. Sharper than the question expected: a verdict is a property an artifact has, not a report about a run.

How does delivery work before self-hosting? — It mostly stops being a question. A reconciler reads source and writes artifacts; where those live is a binding, external first and internal later. The transition looked hard because a pipeline's stages name their targets.

What this effort was right about

Its first question — what is a deployed state, and how does the mesh know it is in one — was marked "everything follows from this", and everything did. Both records above are answers to it: 0058 makes the applier the reporter, and 0063 makes currency a comparison. The effort put the load-bearing question first.

What is NOT closed by this

ADR 0058 names four costs and one of them is a real risk rather than a trade: a reconciler that cannot reach its target retries forever, and without something that notices, the failure is silence — which is the fault this effort exists to catalogue, reintroduced in a new place. That belongs to observability and it is not designed.

The questions (all answered above)

Question Why it matters
What is a deployed state, and how does the mesh know it is in one? Everything follows from this. If a stage reports transport, "deployed" is a claim nobody checked. A desired-state model with reconciliation gives a different answer from a job-completion model.
Does the coordinator dispatch stages, or converge nodes on a declaration? The current model is a state machine over stages. The alternative is that a node is told what should be true and reports what is. The second makes drift visible; the first cannot see it.
How does a change become a pipeline, reliably? Detection has failed for reasons unrelated to the change, silently.
What produces a verdict, and what is it a verdict about? Ties to the lab (ADR 0016) and to a module carrying its own assertions.
How does delivery work before self-hosting, and across the transition? From research 006: source and artifacts start external and are re-bound to internal providers. The coordinator has to be indifferent to which.
Does the three-silo split survive the artifact/part split? ADR 0058 is cardinality-driven, and research 006 renames the thing the cardinality is about.