Not approved as drafted -- four things came out of checking them against each other, and one was a bug that would have broken every upgrade. The bug: 0059 specified Restart=on-failure while 0057 has the host restart onto a new binary by exiting CLEANLY. on-failure does not restart a process that exited zero, so every upgraded node would have been left stopped, having successfully upgraded. Found by reading the two records against each other rather than by either alone. Now Restart=always in all three places that mention it. The host cannot run in a container, and the reason is decisive rather than stylistic: step 0 of the substrate bootstrap installs the container runtime, so a host inside a container would need the thing it exists to install. It would also break 0041 -- copy it onto a machine and run it stops being true when the machine must already have a runtime. Everything above tier 0 is a container; the host is not. That split is the tier boundary, not an inconsistency. systemd is named rather than abstracted. An init is not a dependency in 0041's sense: 0041 is about what must be installed before the host works, and an init is not installed, it is what the machine already is. The unit file is the only systemd-specific artefact and it belongs to the package, so a machine with a different supervisor ships a different package. The mesh is a watchdog, and my first draft was half an answer. Recovery must be local -- nothing dials a node, and a host that cannot start cannot report. But detection is the mesh's, and a local supervisor structurally cannot do it: it sees one process failing and cannot tell a broken machine from a broken release. Only something watching every node can, and that distinction decides whether the response is "fix this machine" or "stop shipping this version". So a host rollout is staged -- a few nodes, wait for heartbeats, continue or stop on silence. Local rollback still needed, because the canary nodes break and because a node offline during the rollout gets the declaration later with no batch around it. The first declaration is the overlay and nothing else. Forced, because a node's address and peers are assigned rather than chosen. But also the way back in: a node reachable over the overlay can be fixed by hand if a later declaration breaks it, and a large first declaration risks a node that is broken and unreachable at once. Also stated plainly, because it reads as a contradiction: nodes reach each other over the overlay and every node consumes from the broker; what 0039 forbids is an inbound CONTROL surface, not reachability. And in 06: no node holds a credential to any control-plane store, for reads or writes. Four ADRs already say this separately and none of them said it in one place. Nodes state over the broker; the owning context writes. With a note that most high-frequency writes are observability's, not the registry's -- routing logs into the registry would be the shared-schema mistake arriving through a door marked performance.
7.0 KiB
status, date, deciders, reconstructed, extends
| status | date | deciders | reconstructed | extends |
|---|---|---|---|---|
| accepted | 2026-08-27 | jochen | false | 0014-build-publish-and-deploy-are-three-silos.md |
58. Delivery ends in a declaration, not in a push to a node
Context
Today a push produces a pipeline with three silos, and the third — deploy — runs once per
module per node, sending a command to every assigned node telling it to install, configure,
start and verify (00-as-is/04).
That third silo is where the as-is document records the most damage:
- "A green pipeline proves transport, not effect." The stages report a message was dispatched and accepted, not that anything is running.
- A service reported started when the container command merely returned. An image pull failure that did not fail the deploy. A package install that 404ed from every mirror while the job went green (04-ISSUES/001).
- A node left on old code after a failed download, with a version marker that had already advanced.
- A verify stage built to close the gap, and never scheduled, because the coordinator's stage list did not include it.
Meanwhile ADR 0037 has given every node a component that does exactly what deploy does — applies state, reads back, reports — and does it continuously rather than once per pipeline. Two mechanisms now change a node, and only one of them checks its work.
Decision
A pipeline ends when the declaration is updated. The node applies it.
The three silos become:
| Silo | Runs | Ends with |
|---|---|---|
| build | once per module | a self-contained artifact |
| publish | once per module | that artifact addressable — an image by digest, a package in the mesh's repository |
| deploy | once, not once per node | the affected nodes' declarations updated to name the new version |
Deploy stops sending commands to nodes. It changes what the control plane says each node should be, which is one write. What happens on the machines is the host's ordinary reconcile, on whatever schedule each node is on.
Why this fixes the failure class rather than patching it
The as-is faults share one shape: the thing that reported success was not the thing that did the work. A coordinator dispatching a command can only report on dispatch.
Under this decision the reporter is the applier. The host already refuses to record a resource until it read it back (ADR 0035), and already fails the whole apply on one failed step (ADR 0008). A package that 404s cannot go green, because nothing between the package manager and the report has an opportunity to be optimistic.
The verify stage disappears as a stage, which is the strongest evidence for this shape: verification stops being a step that can be omitted from a list, and becomes a property of applying at all.
What a pipeline result now means
The honest answer, and it is different from today's:
The declaration is updated, and here is which nodes have applied it.
A pipeline does not wait for every node, because a node may be legitimately switched off for a week (ADR 0036) and a delivery mechanism that blocks on a sleeping laptop is one nobody will use. It reports what landed and what has not landed yet:
delivered declaration updated for 5 nodes
applied 3 of 5
outstanding 2 — last seen 4 days ago, 20 minutes ago
Outstanding is not failure, and conflating them is how the old system got a stall with no error anywhere. A node that has not applied yet is a fact with a timestamp, and it resolves itself when the node comes back.
The host is delivered the same way as everything else
The host is tier 0, which makes it tempting to treat as special. It is not:
- a push to
mesh-hostbuilds a binary; - publish packages it and puts it in the mesh's own package repository — which is a directory of files behind the object store and the proxy, so it needs no new machinery;
- deploy updates each node's declaration to name the new version;
- the new declaration is pushed to each node, and the host applies
package: nox-mesh-hoston arrival, exactly as it applies any other package.
Step 4 is a push and not a poll. The link is already open (ADR 0001), so a node learns of a change when it is made — a node that is offline learns on reconnect, which is what makes outstanding a real category rather than a euphemism for lost.
No new resource type is needed, which is the test of whether this is uniform or a special
case wearing a uniform. The repository is reachable because a file resource put its address in
the package manager's configuration — an ordinary declaration, applied by the same host.
Consequences
- The most consequential boundary in the mesh gets smaller. The as-is calls the fan-out "the most consequential boundary in the mesh" and documents a defect class from the build node having passed through two silos while others had not. There is no fan-out: deploy is one write, and the asymmetry it created cannot arise.
- Detection stays the fragile input, and this does not fix it. A merge that created no pipeline, and nothing said so is upstream of everything here and is untouched.
- Rollback becomes a declaration change, which is a real gain — the previous version is still named in the previous declaration — but nothing here designs how a previous declaration is retained or chosen.
- "Deployed" needs redefining wherever it is used, because it now means told, and the useful fact is applied on node X at time T. Anything reporting deployment state has to move to the second, or it will report success for work that has not happened — the exact fault this record is closing, reintroduced at the reporting layer.
- A node offline for a long time applies a large jump at once, having missed intermediate versions. That is correct — the declaration is a desired state, not a queue of changes — but a machine returning after months applies a very different declaration than it left with, and nothing tests that path.
- The pipeline stops being able to lie and starts being able to be incomplete. That is a better failure mode and it is still a failure mode: a result that is honest about two outstanding nodes is only useful if somebody looks at it.
References
- ADR 0014 — the silos this keeps and redefines the third of.
- ADR 0037 — the applier that makes this possible.
- ADR 0008, ADR 0035 — why the host cannot report success it did not verify.
00-as-is/04— the faults this addresses.