Two corrections and one new decision, all from Jochen catching things. Pushed, not polled. I described updates as landing "on the next reconcile", which reads as polling and is not the design. A declaration arrives as a message on a link that is already open; the host applies it then. Polling over an existing connection would be slower to land AND constant traffic to learn nothing. The timer is for drift and nothing else, and it cannot be replaced by an event for a definitional reason: drift is change the mesh did not make -- somebody edited a managed file, a distribution upgrade replaced a config -- so nothing will ever publish a message about it. Only looking finds it. Separated the heartbeat from the reconcile timer, which I had been conflating. They point in opposite directions and answer different questions: the timer looks at the machine and asks whether it still matches; the heartbeat reports upward and is what makes silence mean something. A node with nothing to do sends nothing, and without a heartbeat that is indistinguishable from a node that stopped. 0059 -- a host that cannot start is rolled back by the service manager. I had left this open on the grounds that recovery meant the host judging its own health. That objection does not survive being asked properly: a keepalive is something else judging the host. The watchdog must be local, because nothing dials a node and a host that cannot start cannot report -- so it is the service manager, which is already there. The failure it prevents is sharper than "the node is down": a host that will not start looks exactly like a machine somebody switched off, which is the one condition this design has deliberately decided not to alarm on. So a bad release reaches every node, each goes quiet, and the mesh reports a fleet of sleeping laptops. Confirmed means started and completed one reconcile -- deliberately not "the link is up", or a laptop on a train would roll itself back. The rollback is a script shipped by the package, not a host subcommand, because a binary that will not start cannot be its own recovery. It rolls back once: a second failure means the machine is the problem, not the binary. Also refined the records checker, which produced a false positive: a proposed record may extend another proposed one, because decisions are drafted in chains and the alternative is marking things accepted to satisfy a check. An accepted document resting on a proposed record still fails, and that was verified. 0057, 0058 and 0059 are all proposed.
136 lines
7.0 KiB
Markdown
136 lines
7.0 KiB
Markdown
---
|
|
status: proposed
|
|
date: 2026-08-27
|
|
deciders: jochen
|
|
reconstructed: false
|
|
extends: 0014-build-publish-and-deploy-are-three-silos.md
|
|
---
|
|
|
|
# 58. Delivery ends in a declaration, not in a push to a node
|
|
|
|
## Context
|
|
|
|
Today a push produces a pipeline with three silos, and the third — **deploy** — runs *once per
|
|
module per node*, sending a command to every assigned node telling it to install, configure,
|
|
start and verify ([`00-as-is/04`](../03-DESIGN/00-as-is/04-delivery.md)).
|
|
|
|
That third silo is where the as-is document records the most damage:
|
|
|
|
- *"A green pipeline proves transport, not effect."* The stages report a message was dispatched
|
|
and accepted, not that anything is running.
|
|
- A service reported started when the container command merely returned. An image pull failure
|
|
that did not fail the deploy. **A package install that 404ed from every mirror while the job
|
|
went green** ([04-ISSUES/001](../04-ISSUES/001-failed-package-install-reports-success/00-report.md)).
|
|
- A node left on old code after a failed download, with a version marker that had already
|
|
advanced.
|
|
- A verify stage built to close the gap, and *never scheduled*, because the coordinator's stage
|
|
list did not include it.
|
|
|
|
Meanwhile [ADR 0037](0037-the-host-applies-it-does-not-decide.md) has given every node a
|
|
component that does exactly what deploy does — applies state, reads back, reports — and does it
|
|
continuously rather than once per pipeline. **Two mechanisms now change a node**, and only one
|
|
of them checks its work.
|
|
|
|
## Decision
|
|
|
|
**A pipeline ends when the declaration is updated. The node applies it.**
|
|
|
|
The three silos become:
|
|
|
|
| Silo | Runs | Ends with |
|
|
|---|---|---|
|
|
| **build** | once per module | a self-contained artifact |
|
|
| **publish** | once per module | that artifact addressable — an image by digest, a package in the mesh's repository |
|
|
| **deploy** | **once, not once per node** | the affected nodes' **declarations updated** to name the new version |
|
|
|
|
**Deploy stops sending commands to nodes.** It changes what the control plane says each node
|
|
should be, which is one write. What happens on the machines is the host's ordinary reconcile,
|
|
on whatever schedule each node is on.
|
|
|
|
### Why this fixes the failure class rather than patching it
|
|
|
|
The as-is faults share one shape: **the thing that reported success was not the thing that did
|
|
the work.** A coordinator dispatching a command can only report on dispatch.
|
|
|
|
Under this decision the reporter *is* the applier. The host already refuses to record a resource
|
|
until it read it back ([ADR 0035](0035-a-picture-is-read-from-what-runs.md)), and already fails
|
|
the whole apply on one failed step ([ADR 0008](0008-a-failed-step-fails-the-job.md)). A package
|
|
that 404s cannot go green, because nothing between the package manager and the report has an
|
|
opportunity to be optimistic.
|
|
|
|
**The verify stage disappears as a stage**, which is the strongest evidence for this shape:
|
|
verification stops being a step that can be omitted from a list, and becomes a property of
|
|
applying at all.
|
|
|
|
### What a pipeline result now means
|
|
|
|
The honest answer, and it is different from today's:
|
|
|
|
> **The declaration is updated, and here is which nodes have applied it.**
|
|
|
|
A pipeline **does not wait for every node**, because a node may be legitimately switched off for
|
|
a week ([ADR 0036](0036-a-node-is-a-managed-machine.md)) and a delivery mechanism that blocks on
|
|
a sleeping laptop is one nobody will use. It reports what landed and what has not landed *yet*:
|
|
|
|
```
|
|
delivered declaration updated for 5 nodes
|
|
applied 3 of 5
|
|
outstanding 2 — last seen 4 days ago, 20 minutes ago
|
|
```
|
|
|
|
**Outstanding is not failure**, and conflating them is how the old system got a stall with no
|
|
error anywhere. A node that has not applied yet is a fact with a timestamp, and it resolves
|
|
itself when the node comes back.
|
|
|
|
### The host is delivered the same way as everything else
|
|
|
|
The host is tier 0, which makes it tempting to treat as special. It is not:
|
|
|
|
1. a push to `mesh-host` builds a binary;
|
|
2. publish packages it and puts it in **the mesh's own package repository** — which is a
|
|
directory of files behind the object store and the proxy, so it needs no new machinery;
|
|
3. deploy updates each node's declaration to name the new version;
|
|
4. **the new declaration is pushed to each node**, and the host applies
|
|
`package: nox-mesh-host` on arrival, exactly as it applies any other package.
|
|
|
|
Step 4 is a push and not a poll. The link is already open
|
|
([ADR 0001](0001-nodes-communicate-over-a-broker.md)), so a node learns of a change when it is
|
|
made — a node that is offline learns on reconnect, which is what makes *outstanding* a real
|
|
category rather than a euphemism for lost.
|
|
|
|
**No new resource type is needed**, which is the test of whether this is uniform or a special
|
|
case wearing a uniform. The repository is reachable because a `file` resource put its address in
|
|
the package manager's configuration — an ordinary declaration, applied by the same host.
|
|
|
|
## Consequences
|
|
|
|
- **The most consequential boundary in the mesh gets smaller.** The as-is calls the fan-out
|
|
*"the most consequential boundary in the mesh"* and documents a defect class from the build
|
|
node having passed through two silos while others had not. **There is no fan-out**: deploy is
|
|
one write, and the asymmetry it created cannot arise.
|
|
- **Detection stays the fragile input, and this does not fix it.** *A merge that created no
|
|
pipeline, and nothing said so* is upstream of everything here and is untouched.
|
|
- **Rollback becomes a declaration change**, which is a real gain — the previous version is
|
|
still named in the previous declaration — but nothing here designs how a previous declaration
|
|
is retained or chosen.
|
|
- **"Deployed" needs redefining wherever it is used**, because it now means *told*, and the
|
|
useful fact is *applied on node X at time T*. Anything reporting deployment state has to move
|
|
to the second, or it will report success for work that has not happened — the exact fault this
|
|
record is closing, reintroduced at the reporting layer.
|
|
- **A node offline for a long time applies a large jump at once**, having missed intermediate
|
|
versions. That is correct — the declaration is a desired state, not a queue of changes — but a
|
|
machine returning after months applies a very different declaration than it left with, and
|
|
nothing tests that path.
|
|
- **The pipeline stops being able to lie and starts being able to be incomplete.** That is a
|
|
better failure mode and it is still a failure mode: a result that is honest about two
|
|
outstanding nodes is only useful if somebody looks at it.
|
|
|
|
## References
|
|
|
|
- [ADR 0014](0014-build-publish-and-deploy-are-three-silos.md) — the silos this keeps and
|
|
redefines the third of.
|
|
- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — the applier that makes this possible.
|
|
- [ADR 0008](0008-a-failed-step-fails-the-job.md), [ADR 0035](0035-a-picture-is-read-from-what-runs.md) —
|
|
why the host cannot report success it did not verify.
|
|
- [`00-as-is/04`](../03-DESIGN/00-as-is/04-delivery.md) — the faults this addresses.
|