Files
hq/02-DECISIONS/0010-delivery.md
T
jschoubben b4607dfc03 Numbers are identity; the reading order is a generated, checked index
Decided after measuring what renumbering actually costs: 96 references in code
comments across two repositories, none of which would have failed to compile.
They would have pointed at the wrong reasoning, which is worse than a broken
link because nothing reports it.

So a number identifies a record and never changes. It cannot also be a
position -- a position moves when the set changes, and an identity that moves
is not one.

The reading order moves into an index generated from each record's `topic:`.
Six topics, in the order somebody learns the system.

The index is WRITTEN rather than only generated on demand, which reverses what
this repository previously said. The reason it said otherwise is that a
hand-written index drifts -- but a reader looking at the folder on a forge sees
the folder, not a command, and the drift objection is answered by checking
rather than by refusing to write one. That is §5's own rule: a rule states how
it is checked.

Two checks, both confirmed to bite. index.py fails when the written order no
longer matches the records. records.py fails when a record has no topic or one
nobody defined -- the quiet failure being a record that vanishes from the order
rather than appearing in the wrong place.
2026-08-28 23:39:18 +02:00

6.7 KiB

topic, status, date, deciders, reconstructed
topic status date deciders reconstructed
what runs on it accepted 2026-08-28 jochen false

10. Delivery

Consolidated 2026-08-28 from five records.

Delivery is a comparison, not a pipeline

The control plane holds what source exists and what has been built from it, and builds the difference. A change becomes a build because source is ahead of artifacts — answerable at any moment — rather than because a message arrived.

An event makes it fast. Nothing makes it necessary. A missed notification costs latency and cannot cost correctness.

That is the same shape the host uses on a machine, one layer up:

reconciles against
the control plane artifacts source
the host machine state declarations

This is not the current coordinator repaired. That is a state machine over stages; the value of it here is as a catalogue of the ways this fails, and it has been used for exactly that.

What disappears is the pipeline as a state machine — no stage list something can be omitted from, which is how a verify stage was built and never scheduled, and no run to lose.

Currency is the whole input closure

An artifact is out of date when its source moved, or anything it was built against moved. So a shared library changing invalidates everything with a transitive build edge to it, in dependency order, because a module cannot be built against a new library until it exists.

The module graph is a prerequisite of this, not an enabler of it. Without it there is no rebuild set and no ordering, and this cannot be implemented.

An artifact is build output, never a source tree

Compiled and bundled with its dependency graph inlined. A deploy is extract-and-run and touches no network.

The consequence is the whole cost of the decision: anything not in the build output does not ship. Every file kind had to be brought into that rule separately, and each was discovered by something silently not happening after a deploy — migrations reading a source layout, provisioning scripts reading a source layout, selection files never packaged at all.

Three silos, and the third is not a stage

The cardinality observation holds and is what the split is for:

silo runs ends with
build once per module a self-contained artifact
publish once per module that artifact addressable — an image by digest, a package in the mesh's repository
deploy once, not once per node the affected nodes' declarations updated

Deploy stops sending commands to nodes. It changes what the control plane says each node should be, which is one write. What happens on the machines is the host's ordinary reconcile.

Why this fixes the failure class rather than patching it. Every recorded fault shares one shape: the thing that reported success was not the thing that did the work. A coordinator dispatching a command can only report on dispatch. Under this the reporter is the applier — which already refuses to record a resource until it read it back, and already fails the whole apply on one failed step.

The verify stage disappears as a stage, which is the strongest evidence for the shape: verification stops being a step that can be omitted from a list and becomes a property of applying at all.

There is no fan-out, so the defect class that came from the build node having passed through two silos while others had not cannot arise.

A step that fails must fail the job

A step that fails and lets the job continue reports success for work that did not happen. Absence of an error is not evidence of an effect.

This is the mesh's most consistent failure shape, and it is not incidental — it is what stage reporting measured. Documented instances: a service reported started when the container command merely returned; an image pull failure that did not fail the deploy; a package install that 404'd from every mirror while the job went green; a node left on old code after a failed download with a version marker that had already advanced.

The verdict is tiered

An artifact may not be declared until something has judged it fit. Two tiers, because one gate would be both slow and unreliable:

judged by when
the module's own tests the build always — this is most of it
the lab a raised scenario when an assertion genuinely needs a mesh

A lab scenario takes tens of seconds and can fail for reasons that have nothing to do with the artifact, and a shared-library change produces a cascade of dozens. One expensive non-deterministic gate fails in both directions: a flaky run marks a good artifact unfit, a lucky one marks a bad artifact fit, and neither failure looks like itself.

A run that failed environmentally is not a verdict. A machine that would not boot says nothing about the artifact, and recording it as unfit is the same untruth as recording a dispatch as a deploy.

What a result means

The declaration is updated, and here is which nodes have applied it.

A pipeline does not wait for every node — one may be legitimately switched off for a week, and a delivery mechanism that blocks on a sleeping laptop is one nobody will use.

delivered      declaration updated for 5 nodes
applied        3 of 5
outstanding    2 — last seen 4 days ago, 20 minutes ago

Outstanding is not failure, and conflating them is how the old system produced a stall with no error anywhere.

What must exist first

  1. The module graph, with build edges. No graph, no rebuild set and no ordering.
  2. A recorded input closure per artifact, so currency is answerable without building.
  3. Something that notices a reconciler is not converging. Below.

Open, and the first is the real risk

  • A loop that will not converge is harder to debug than a job that failed. A failed job stops and names its step; a reconciler retries forever. Without something that notices this has been trying for an hour, the failure is silence — the fault this removes, reintroduced in a new place.
  • The run identity people use is lost. Did my change go out? is answerable today by opening a pipeline. Something must replace that or this is worse to live with, whatever its properties.
  • Does a fit artifact declare itself? If it does, merging to main deploys to production — which may be wanted and is far too large a property to acquire by omission.
  • Rebuild storms are mostly behaviourally empty. Reproducible builds would stop a cascade at the first module whose output did not move; without them one commit redeploys the fleet for no change in behaviour.
  • Detection stays the fragile input for latency, though no longer for correctness.