Resolve the host lifecycle's open items, and say how the host is delivered

The upgrade question turned out to be a delivery question, so 0058 answers
both.

Today's third silo runs once per node and sends each one a command to install
and start. That is where the as-is records a package install that 404ed from
every mirror while the job went green, an image pull failure that did not fail
the deploy, and a verify stage that was built and never scheduled because it
was missing from a list.

The shape underneath all of those is that the thing reporting success was not
the thing doing the work. Meanwhile ADR 0037 has given every node a component
that applies state, reads back and reports -- so two mechanisms now change a
node and only one checks its work.

0058: a pipeline ends when the declaration is updated. Deploy stops sending
commands to nodes and becomes one write. The host applies it on its next
reconcile, and the host cannot report success it did not verify. The verify
stage disappears as a stage, which is the point -- verification stops being a
step that can be left off a list.

A pipeline result now means "the declaration is updated, and here is which
nodes have applied it". It does not wait for every node, because a node may be
legitimately switched off for a week. Outstanding is reported separately from
failed, since conflating them is how the old system produced a stall with no
error anywhere.

The host is delivered by exactly this path and needs no new resource type: a
`file` writes the package manager's config pointing at the mesh's repository, a
`package` names the version. Added a step I had missed -- before exiting for a
restart, the host runs the new binary once. A package can install something
that does not execute here, and that turns "the node never came back" into "the
apply failed and said why".

Six open items resolved: re-enrolment is decided when the token is issued and
revokes the previous identity; the mesh keeps a recovery copy of what each node
reports it owns, which un-strands the orphans; last-contact is reported with no
threshold, because a laptop off for three weeks is doing nothing wrong;
adoption always completes but a failed line makes a node ineligible for
assignment; a briefing is a structured document whose outcome is computed from
its lines; and the token is printed once and carried by hand, which is the
property that makes it worth anything.

Still open and named: automatic rollback of a host version that will not start.

0057 and 0058 are both proposed.
This commit is contained in:
2026-08-27 21:53:38 +02:00
parent 2204b01909
commit aeea2a9f9a
2 changed files with 306 additions and 36 deletions
@@ -0,0 +1,130 @@
---
status: proposed
date: 2026-08-27
deciders: jochen
reconstructed: false
extends: 0014-build-publish-and-deploy-are-three-silos.md
---
# 58. Delivery ends in a declaration, not in a push to a node
## Context
Today a push produces a pipeline with three silos, and the third — **deploy** — runs *once per
module per node*, sending a command to every assigned node telling it to install, configure,
start and verify ([`00-as-is/04`](../03-DESIGN/00-as-is/04-delivery.md)).
That third silo is where the as-is document records the most damage:
- *"A green pipeline proves transport, not effect."* The stages report a message was dispatched
and accepted, not that anything is running.
- A service reported started when the container command merely returned. An image pull failure
that did not fail the deploy. **A package install that 404ed from every mirror while the job
went green** ([04-ISSUES/001](../04-ISSUES/001-failed-package-install-reports-success/00-report.md)).
- A node left on old code after a failed download, with a version marker that had already
advanced.
- A verify stage built to close the gap, and *never scheduled*, because the coordinator's stage
list did not include it.
Meanwhile [ADR 0037](0037-the-host-applies-it-does-not-decide.md) has given every node a
component that does exactly what deploy does — applies state, reads back, reports — and does it
continuously rather than once per pipeline. **Two mechanisms now change a node**, and only one
of them checks its work.
## Decision
**A pipeline ends when the declaration is updated. The node applies it.**
The three silos become:
| Silo | Runs | Ends with |
|---|---|---|
| **build** | once per module | a self-contained artifact |
| **publish** | once per module | that artifact addressable — an image by digest, a package in the mesh's repository |
| **deploy** | **once, not once per node** | the affected nodes' **declarations updated** to name the new version |
**Deploy stops sending commands to nodes.** It changes what the control plane says each node
should be, which is one write. What happens on the machines is the host's ordinary reconcile,
on whatever schedule each node is on.
### Why this fixes the failure class rather than patching it
The as-is faults share one shape: **the thing that reported success was not the thing that did
the work.** A coordinator dispatching a command can only report on dispatch.
Under this decision the reporter *is* the applier. The host already refuses to record a resource
until it read it back ([ADR 0035](0035-a-picture-is-read-from-what-runs.md)), and already fails
the whole apply on one failed step ([ADR 0008](0008-a-failed-step-fails-the-job.md)). A package
that 404s cannot go green, because nothing between the package manager and the report has an
opportunity to be optimistic.
**The verify stage disappears as a stage**, which is the strongest evidence for this shape:
verification stops being a step that can be omitted from a list, and becomes a property of
applying at all.
### What a pipeline result now means
The honest answer, and it is different from today's:
> **The declaration is updated, and here is which nodes have applied it.**
A pipeline **does not wait for every node**, because a node may be legitimately switched off for
a week ([ADR 0036](0036-a-node-is-a-managed-machine.md)) and a delivery mechanism that blocks on
a sleeping laptop is one nobody will use. It reports what landed and what has not landed *yet*:
```
delivered declaration updated for 5 nodes
applied 3 of 5
outstanding 2 — last seen 4 days ago, 20 minutes ago
```
**Outstanding is not failure**, and conflating them is how the old system got a stall with no
error anywhere. A node that has not applied yet is a fact with a timestamp, and it resolves
itself when the node comes back.
### The host is delivered the same way as everything else
The host is tier 0, which makes it tempting to treat as special. It is not:
1. a push to `mesh-host` builds a binary;
2. publish packages it and puts it in **the mesh's own package repository** — which is a
directory of files behind the object store and the proxy, so it needs no new machinery;
3. deploy updates each node's declaration to name the new version;
4. each host applies `package: nox-mesh-host` on its next reconcile, exactly as it applies any
other package.
**No new resource type is needed**, which is the test of whether this is uniform or a special
case wearing a uniform. The repository is reachable because a `file` resource put its address in
the package manager's configuration — an ordinary declaration, applied by the same host.
## Consequences
- **The most consequential boundary in the mesh gets smaller.** The as-is calls the fan-out
*"the most consequential boundary in the mesh"* and documents a defect class from the build
node having passed through two silos while others had not. **There is no fan-out**: deploy is
one write, and the asymmetry it created cannot arise.
- **Detection stays the fragile input, and this does not fix it.** *A merge that created no
pipeline, and nothing said so* is upstream of everything here and is untouched.
- **Rollback becomes a declaration change**, which is a real gain — the previous version is
still named in the previous declaration — but nothing here designs how a previous declaration
is retained or chosen.
- **"Deployed" needs redefining wherever it is used**, because it now means *told*, and the
useful fact is *applied on node X at time T*. Anything reporting deployment state has to move
to the second, or it will report success for work that has not happened — the exact fault this
record is closing, reintroduced at the reporting layer.
- **A node offline for a long time applies a large jump at once**, having missed intermediate
versions. That is correct — the declaration is a desired state, not a queue of changes — but a
machine returning after months applies a very different declaration than it left with, and
nothing tests that path.
- **The pipeline stops being able to lie and starts being able to be incomplete.** That is a
better failure mode and it is still a failure mode: a result that is honest about two
outstanding nodes is only useful if somebody looks at it.
## References
- [ADR 0014](0014-build-publish-and-deploy-are-three-silos.md) — the silos this keeps and
redefines the third of.
- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — the applier that makes this possible.
- [ADR 0008](0008-a-failed-step-fails-the-job.md), [ADR 0035](0035-a-picture-is-read-from-what-runs.md) —
why the host cannot report success it did not verify.
- [`00-as-is/04`](../03-DESIGN/00-as-is/04-delivery.md) — the faults this addresses.