Files
hq/02-DECISIONS/0023-delivery.md
T
jschoubben e1febe8e0f Renumber the records 1 to 23
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18,
19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only
the archaeology of what used to be there.

Renumbered contiguously. Renames run in ascending order, so every target number
is already free and no two files ever collide.

The reference rewrite is one simultaneous pass rather than a sequence of
replacements. Numbers moved into slots other numbers were vacating -- the node
host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time
would have cascaded and silently pointed things at the wrong record.

Seven plain-text references survived the merges as prose rather than links,
naming records that no longer existed: the enrolment token, the link boundary,
what a declaration is, reachability, the repository structure. Each mapped to
the consolidated record that now holds it.

Verified rather than assumed: every [ADR NNNN](path) link now has matching text
and target, checked across the whole repository, and the checker passes.

Frontmatter `consolidates:` lists dropped -- they named records that are gone,
and each consolidated record already says in prose what it absorbed.
2026-08-28 23:28:34 +02:00

145 lines
6.7 KiB
Markdown

---
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 23. Delivery
*Consolidated 2026-08-28 from five records.*
## Delivery is a comparison, not a pipeline
**The control plane holds what source exists and what has been built from it, and builds the
difference.** A change becomes a build because source is **ahead of artifacts** — answerable at
any moment — rather than because a message arrived.
**An event makes it fast. Nothing makes it necessary.** A missed notification costs latency and
cannot cost correctness.
That is the same shape the host uses on a machine, one layer up:
| | reconciles | against |
|---|---|---|
| the control plane | artifacts | source |
| the host | machine state | declarations |
**This is not the current coordinator repaired.** That is a state machine over stages; the value
of it here is as a catalogue of the ways this fails, and it has been used for exactly that.
**What disappears is the pipeline as a state machine** — no stage list something can be omitted
from, which is how a verify stage was built and never scheduled, and no run to lose.
### Currency is the whole input closure
**An artifact is out of date when its source moved, or anything it was built against moved.** So a
shared library changing invalidates everything with a transitive build edge to it, in dependency
order, because a module cannot be built against a new library until it exists.
**The module graph is a prerequisite of this, not an enabler of it.** Without it there is no
rebuild set and no ordering, and this cannot be implemented.
## An artifact is build output, never a source tree
Compiled and bundled with its dependency graph inlined. **A deploy is extract-and-run and touches
no network.**
The consequence is the whole cost of the decision: **anything not in the build output does not
ship.** Every file kind had to be brought into that rule separately, and each was discovered by
something silently not happening after a deploy — migrations reading a source layout,
provisioning scripts reading a source layout, selection files never packaged at all.
## Three silos, and the third is not a stage
The cardinality observation holds and is what the split is for:
| silo | runs | ends with |
|---|---|---|
| **build** | once per module | a self-contained artifact |
| **publish** | once per module | that artifact addressable — an image by digest, a package in the mesh's repository |
| **deploy** | **once, not once per node** | the affected nodes' **declarations updated** |
**Deploy stops sending commands to nodes.** It changes what the control plane says each node
should be, which is one write. What happens on the machines is the host's ordinary reconcile.
**Why this fixes the failure class rather than patching it.** Every recorded fault shares one
shape: *the thing that reported success was not the thing that did the work.* A coordinator
dispatching a command can only report on dispatch. Under this the reporter **is** the applier —
which already refuses to record a resource until it read it back, and already fails the whole
apply on one failed step.
**The verify stage disappears as a stage**, which is the strongest evidence for the shape:
verification stops being a step that can be omitted from a list and becomes a property of applying
at all.
**There is no fan-out**, so the defect class that came from the build node having passed through
two silos while others had not cannot arise.
## A step that fails must fail the job
A step that fails and lets the job continue **reports success for work that did not happen**.
Absence of an error is not evidence of an effect.
This is the mesh's most consistent failure shape, and it is not incidental — it is what stage
reporting measured. Documented instances: a service reported started when the container command
merely returned; an image pull failure that did not fail the deploy; a package install that 404'd
from every mirror while the job went green; a node left on old code after a failed download with
a version marker that had already advanced.
## The verdict is tiered
An artifact may not be declared until something has judged it fit. **Two tiers, because one gate
would be both slow and unreliable:**
| | judged by | when |
|---|---|---|
| the module's own tests | the build | **always** — this is most of it |
| the lab | a raised scenario | when an assertion genuinely needs a mesh |
A lab scenario takes tens of seconds and can fail for reasons that have nothing to do with the
artifact, and a shared-library change produces a cascade of dozens. One expensive
non-deterministic gate fails in both directions: a flaky run marks a good artifact unfit, a lucky
one marks a bad artifact fit, and **neither failure looks like itself.**
**A run that failed environmentally is not a verdict.** A machine that would not boot says nothing
about the artifact, and recording it as *unfit* is the same untruth as recording a dispatch as a
deploy.
## What a result means
> **The declaration is updated, and here is which nodes have applied it.**
A pipeline does not wait for every node — one may be legitimately switched off for a week, and a
delivery mechanism that blocks on a sleeping laptop is one nobody will use.
```
delivered declaration updated for 5 nodes
applied 3 of 5
outstanding 2 — last seen 4 days ago, 20 minutes ago
```
**Outstanding is not failure**, and conflating them is how the old system produced a stall with no
error anywhere.
## What must exist first
1. **The module graph, with build edges.** No graph, no rebuild set and no ordering.
2. **A recorded input closure per artifact**, so currency is answerable without building.
3. **Something that notices a reconciler is not converging.** Below.
## Open, and the first is the real risk
- **A loop that will not converge is harder to debug than a job that failed.** A failed job stops
and names its step; a reconciler retries forever. Without something that notices *this has been
trying for an hour*, the failure is **silence** — the fault this removes, reintroduced in a new
place.
- **The run identity people use is lost.** *Did my change go out?* is answerable today by opening
a pipeline. Something must replace that or this is worse to live with, whatever its properties.
- **Does a fit artifact declare itself?** If it does, merging to main deploys to production —
which may be wanted and is far too large a property to acquire by omission.
- **Rebuild storms are mostly behaviourally empty.** Reproducible builds would stop a cascade at
the first module whose output did not move; without them one commit redeploys the fleet for no
change in behaviour.
- **Detection stays the fragile input for latency**, though no longer for correctness.