Files
hq/02-DECISIONS/0010-delivery.md
jschoubben 918dc04916 What this actually is, and three things that were assumed
Four things settled by talking them through, all of which had been true in
somebody's head and written nowhere.

It is not a mesh in the peer-to-peer sense and will not become one. 0001 now
says what it is instead: machines linked by a private network, one node holding
knowledge of all of them, modules as the way anything is built and delivered,
and agents hired onto nodes to do the work. The word describes what machines
can reach, not how they are governed. "Master" overstates it the other way --
nothing needs that node to keep running, only to change.

0006 gains the option that would make it a real mesh, recorded as considered
rather than rejected by silence: every node holding the whole inventory, a
replication process, an elected master with promotion on failure. What settles
it is not the complexity but that it still would not deliver the name, because
application databases are not replicated -- so a genuine peer-to-peer mesh
means becoming a replicated database system for every consumer's data too. That
is a larger product than the thing it would support.

Also in 0006: three central roles, not one. Losing the control plane costs
change, losing the broker costs being told anything, and losing the hub costs
nodes in different places reaching each other at all -- which is operation, not
administration. Whether they are one node is not decided.

And SSH access is identity's. It appeared three times as something that uses
the overlay and never as something the mesh provides, which reads as settled
when nothing decided it. Nobody else could: the mesh is the only thing that
knows which humans and agents exist and which nodes they may reach. Node to
node SSH stays out -- the host has no inbound control surface by decision, and
nodes reaching each other that way is a second control path through the back
door.

0007 gains the requirement underneath all of it. Reachability was recorded as a
fact to track and never as a thing some node must have. The broker's node and
the hub must be dialable by every node at a stable address, or nothing can join
and a disconnected node cannot return. A mesh entirely behind NAT cannot be
raised. That is a precondition and it belongs with the others.

The link staying on the underlay is also argued now rather than asserted. At
join time it is forced; afterwards it is a choice, and the reason is that a
repair channel carried over the thing being repaired is not one. Moving it onto
the overlay, with fallback, is recorded as open with what it would have to get
right -- a WireGuard interface has no link state to test, and a silent fallback
is this repository's recurring fault in a new place.

0010 says in one line what was the intention throughout: the module system is
the CI/CD. Not a pipeline beside the mesh. Build, test, publish and deploy are
one reconciliation seen at four points, which is why a thing that cannot be a
module cannot be delivered.
2026-08-29 13:01:36 +02:00

7.7 KiB

topic, status, date, deciders, reconstructed
topic status date deciders reconstructed
what runs on it accepted 2026-08-28 jochen false

10. Delivery

Consolidated 2026-08-28 from five records.

Delivery is a comparison, not a pipeline

The control plane holds what source exists and what has been built from it, and builds the difference. A change becomes a build because source is ahead of artifacts — answerable at any moment — rather than because a message arrived.

An event makes it fast. Nothing makes it necessary. A missed notification costs latency and cannot cost correctness.

That is the same shape the host uses on a machine, one layer up:

reconciles against
the control plane artifacts source
the host machine state declarations

This is not the current coordinator repaired. That is a state machine over stages; the value of it here is as a catalogue of the ways this fails, and it has been used for exactly that.

The module system is the CI/CD

Written 2026-08-29, because this was the intention throughout and was never stated in one line.

There is no pipeline product beside the mesh, and there is not going to be one. A module declares what it is (ADR 0009); the control plane notices its source is ahead of its artifacts and builds it; the graph says what else that invalidates; the node that should run it is told. Build, test, publish and deploy are the same reconciliation seen at four points, not four stages wired together.

Which is why the module system is the core of the setup rather than one component of it. Every other layer is carried by it: the substrate is modules the bundle raises before there is a mesh, the control plane is a module, and an application is a module with a different manifest. A thing that cannot be expressed as a module cannot be delivered at all — that is a real constraint, and it is the one keeping a second delivery mechanism from growing beside this one.

What disappears is the pipeline as a state machine — no stage list something can be omitted from, which is how a verify stage was built and never scheduled, and no run to lose.

Currency is the whole input closure

An artifact is out of date when its source moved, or anything it was built against moved. So a shared library changing invalidates everything with a transitive build edge to it, in dependency order, because a module cannot be built against a new library until it exists.

The module graph is a prerequisite of this, not an enabler of it. Without it there is no rebuild set and no ordering, and this cannot be implemented.

An artifact is build output, never a source tree

Compiled and bundled with its dependency graph inlined. A deploy is extract-and-run and touches no network.

The consequence is the whole cost of the decision: anything not in the build output does not ship. Every file kind had to be brought into that rule separately, and each was discovered by something silently not happening after a deploy — migrations reading a source layout, provisioning scripts reading a source layout, selection files never packaged at all.

Three silos, and the third is not a stage

The cardinality observation holds and is what the split is for:

silo runs ends with
build once per module a self-contained artifact
publish once per module that artifact addressable — an image by digest, a package in the mesh's repository
deploy once, not once per node the affected nodes' declarations updated

Deploy stops sending commands to nodes. It changes what the control plane says each node should be, which is one write. What happens on the machines is the host's ordinary reconcile.

Why this fixes the failure class rather than patching it. Every recorded fault shares one shape: the thing that reported success was not the thing that did the work. A coordinator dispatching a command can only report on dispatch. Under this the reporter is the applier — which already refuses to record a resource until it read it back, and already fails the whole apply on one failed step.

The verify stage disappears as a stage, which is the strongest evidence for the shape: verification stops being a step that can be omitted from a list and becomes a property of applying at all.

There is no fan-out, so the defect class that came from the build node having passed through two silos while others had not cannot arise.

A step that fails must fail the job

A step that fails and lets the job continue reports success for work that did not happen. Absence of an error is not evidence of an effect.

This is the mesh's most consistent failure shape, and it is not incidental — it is what stage reporting measured. Documented instances: a service reported started when the container command merely returned; an image pull failure that did not fail the deploy; a package install that 404'd from every mirror while the job went green; a node left on old code after a failed download with a version marker that had already advanced.

The verdict is tiered

An artifact may not be declared until something has judged it fit. Two tiers, because one gate would be both slow and unreliable:

judged by when
the module's own tests the build always — this is most of it
the lab a raised scenario when an assertion genuinely needs a mesh

A lab scenario takes tens of seconds and can fail for reasons that have nothing to do with the artifact, and a shared-library change produces a cascade of dozens. One expensive non-deterministic gate fails in both directions: a flaky run marks a good artifact unfit, a lucky one marks a bad artifact fit, and neither failure looks like itself.

A run that failed environmentally is not a verdict. A machine that would not boot says nothing about the artifact, and recording it as unfit is the same untruth as recording a dispatch as a deploy.

What a result means

The declaration is updated, and here is which nodes have applied it.

A pipeline does not wait for every node — one may be legitimately switched off for a week, and a delivery mechanism that blocks on a sleeping laptop is one nobody will use.

delivered      declaration updated for 5 nodes
applied        3 of 5
outstanding    2 — last seen 4 days ago, 20 minutes ago

Outstanding is not failure, and conflating them is how the old system produced a stall with no error anywhere.

What must exist first

  1. The module graph, with build edges. No graph, no rebuild set and no ordering.
  2. A recorded input closure per artifact, so currency is answerable without building.
  3. Something that notices a reconciler is not converging. Below.

Open, and the first is the real risk

  • A loop that will not converge is harder to debug than a job that failed. A failed job stops and names its step; a reconciler retries forever. Without something that notices this has been trying for an hour, the failure is silence — the fault this removes, reintroduced in a new place.
  • The run identity people use is lost. Did my change go out? is answerable today by opening a pipeline. Something must replace that or this is worse to live with, whatever its properties.
  • Does a fit artifact declare itself? If it does, merging to main deploys to production — which may be wanted and is far too large a property to acquire by omission.
  • Rebuild storms are mostly behaviourally empty. Reproducible builds would stop a cascade at the first module whose output did not move; without them one commit redeploys the fleet for no change in behaviour.
  • Detection stays the fragile input for latency, though no longer for correctness.