Order the records the way the system is learned

Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
This commit is contained in:
2026-08-28 23:30:42 +02:00
parent e1febe8e0f
commit 333356cff3
85 changed files with 471 additions and 465 deletions
+144
View File
@@ -0,0 +1,144 @@
---
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 10. Delivery
*Consolidated 2026-08-28 from five records.*
## Delivery is a comparison, not a pipeline
**The control plane holds what source exists and what has been built from it, and builds the
difference.** A change becomes a build because source is **ahead of artifacts** — answerable at
any moment — rather than because a message arrived.
**An event makes it fast. Nothing makes it necessary.** A missed notification costs latency and
cannot cost correctness.
That is the same shape the host uses on a machine, one layer up:
| | reconciles | against |
|---|---|---|
| the control plane | artifacts | source |
| the host | machine state | declarations |
**This is not the current coordinator repaired.** That is a state machine over stages; the value
of it here is as a catalogue of the ways this fails, and it has been used for exactly that.
**What disappears is the pipeline as a state machine** — no stage list something can be omitted
from, which is how a verify stage was built and never scheduled, and no run to lose.
### Currency is the whole input closure
**An artifact is out of date when its source moved, or anything it was built against moved.** So a
shared library changing invalidates everything with a transitive build edge to it, in dependency
order, because a module cannot be built against a new library until it exists.
**The module graph is a prerequisite of this, not an enabler of it.** Without it there is no
rebuild set and no ordering, and this cannot be implemented.
## An artifact is build output, never a source tree
Compiled and bundled with its dependency graph inlined. **A deploy is extract-and-run and touches
no network.**
The consequence is the whole cost of the decision: **anything not in the build output does not
ship.** Every file kind had to be brought into that rule separately, and each was discovered by
something silently not happening after a deploy — migrations reading a source layout,
provisioning scripts reading a source layout, selection files never packaged at all.
## Three silos, and the third is not a stage
The cardinality observation holds and is what the split is for:
| silo | runs | ends with |
|---|---|---|
| **build** | once per module | a self-contained artifact |
| **publish** | once per module | that artifact addressable — an image by digest, a package in the mesh's repository |
| **deploy** | **once, not once per node** | the affected nodes' **declarations updated** |
**Deploy stops sending commands to nodes.** It changes what the control plane says each node
should be, which is one write. What happens on the machines is the host's ordinary reconcile.
**Why this fixes the failure class rather than patching it.** Every recorded fault shares one
shape: *the thing that reported success was not the thing that did the work.* A coordinator
dispatching a command can only report on dispatch. Under this the reporter **is** the applier —
which already refuses to record a resource until it read it back, and already fails the whole
apply on one failed step.
**The verify stage disappears as a stage**, which is the strongest evidence for the shape:
verification stops being a step that can be omitted from a list and becomes a property of applying
at all.
**There is no fan-out**, so the defect class that came from the build node having passed through
two silos while others had not cannot arise.
## A step that fails must fail the job
A step that fails and lets the job continue **reports success for work that did not happen**.
Absence of an error is not evidence of an effect.
This is the mesh's most consistent failure shape, and it is not incidental — it is what stage
reporting measured. Documented instances: a service reported started when the container command
merely returned; an image pull failure that did not fail the deploy; a package install that 404'd
from every mirror while the job went green; a node left on old code after a failed download with
a version marker that had already advanced.
## The verdict is tiered
An artifact may not be declared until something has judged it fit. **Two tiers, because one gate
would be both slow and unreliable:**
| | judged by | when |
|---|---|---|
| the module's own tests | the build | **always** — this is most of it |
| the lab | a raised scenario | when an assertion genuinely needs a mesh |
A lab scenario takes tens of seconds and can fail for reasons that have nothing to do with the
artifact, and a shared-library change produces a cascade of dozens. One expensive
non-deterministic gate fails in both directions: a flaky run marks a good artifact unfit, a lucky
one marks a bad artifact fit, and **neither failure looks like itself.**
**A run that failed environmentally is not a verdict.** A machine that would not boot says nothing
about the artifact, and recording it as *unfit* is the same untruth as recording a dispatch as a
deploy.
## What a result means
> **The declaration is updated, and here is which nodes have applied it.**
A pipeline does not wait for every node — one may be legitimately switched off for a week, and a
delivery mechanism that blocks on a sleeping laptop is one nobody will use.
```
delivered declaration updated for 5 nodes
applied 3 of 5
outstanding 2 — last seen 4 days ago, 20 minutes ago
```
**Outstanding is not failure**, and conflating them is how the old system produced a stall with no
error anywhere.
## What must exist first
1. **The module graph, with build edges.** No graph, no rebuild set and no ordering.
2. **A recorded input closure per artifact**, so currency is answerable without building.
3. **Something that notices a reconciler is not converging.** Below.
## Open, and the first is the real risk
- **A loop that will not converge is harder to debug than a job that failed.** A failed job stops
and names its step; a reconciler retries forever. Without something that notices *this has been
trying for an hour*, the failure is **silence** — the fault this removes, reintroduced in a new
place.
- **The run identity people use is lost.** *Did my change go out?* is answerable today by opening
a pipeline. Something must replace that or this is worse to live with, whatever its properties.
- **Does a fit artifact declare itself?** If it does, merging to main deploys to production —
which may be wanted and is far too large a property to acquire by omission.
- **Rebuild storms are mostly behaviourally empty.** Reproducible builds would stop a cascade at
the first module whose output did not move; without them one commit redeploys the fleet for no
change in behaviour.
- **Detection stays the fragile input for latency**, though no longer for correctness.