Resolve the host lifecycle's open items, and say how the host is delivered
The upgrade question turned out to be a delivery question, so 0058 answers both. Today's third silo runs once per node and sends each one a command to install and start. That is where the as-is records a package install that 404ed from every mirror while the job went green, an image pull failure that did not fail the deploy, and a verify stage that was built and never scheduled because it was missing from a list. The shape underneath all of those is that the thing reporting success was not the thing doing the work. Meanwhile ADR 0037 has given every node a component that applies state, reads back and reports -- so two mechanisms now change a node and only one checks its work. 0058: a pipeline ends when the declaration is updated. Deploy stops sending commands to nodes and becomes one write. The host applies it on its next reconcile, and the host cannot report success it did not verify. The verify stage disappears as a stage, which is the point -- verification stops being a step that can be left off a list. A pipeline result now means "the declaration is updated, and here is which nodes have applied it". It does not wait for every node, because a node may be legitimately switched off for a week. Outstanding is reported separately from failed, since conflating them is how the old system produced a stall with no error anywhere. The host is delivered by exactly this path and needs no new resource type: a `file` writes the package manager's config pointing at the mesh's repository, a `package` names the version. Added a step I had missed -- before exiting for a restart, the host runs the new binary once. A package can install something that does not execute here, and that turns "the node never came back" into "the apply failed and said why". Six open items resolved: re-enrolment is decided when the token is issued and revokes the previous identity; the mesh keeps a recovery copy of what each node reports it owns, which un-strands the orphans; last-contact is reported with no threshold, because a laptop off for three weeks is doing nothing wrong; adoption always completes but a failed line makes a node ineligible for assignment; a briefing is a structured document whose outcome is computed from its lines; and the token is printed once and carried by hand, which is the property that makes it worth anything. Still open and named: automatic rollback of a host version that will not start. 0057 and 0058 are both proposed.
This commit is contained in:
@@ -0,0 +1,130 @@
|
||||
---
|
||||
status: proposed
|
||||
date: 2026-08-27
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 0014-build-publish-and-deploy-are-three-silos.md
|
||||
---
|
||||
|
||||
# 58. Delivery ends in a declaration, not in a push to a node
|
||||
|
||||
## Context
|
||||
|
||||
Today a push produces a pipeline with three silos, and the third — **deploy** — runs *once per
|
||||
module per node*, sending a command to every assigned node telling it to install, configure,
|
||||
start and verify ([`00-as-is/04`](../03-DESIGN/00-as-is/04-delivery.md)).
|
||||
|
||||
That third silo is where the as-is document records the most damage:
|
||||
|
||||
- *"A green pipeline proves transport, not effect."* The stages report a message was dispatched
|
||||
and accepted, not that anything is running.
|
||||
- A service reported started when the container command merely returned. An image pull failure
|
||||
that did not fail the deploy. **A package install that 404ed from every mirror while the job
|
||||
went green** ([04-ISSUES/001](../04-ISSUES/001-failed-package-install-reports-success/00-report.md)).
|
||||
- A node left on old code after a failed download, with a version marker that had already
|
||||
advanced.
|
||||
- A verify stage built to close the gap, and *never scheduled*, because the coordinator's stage
|
||||
list did not include it.
|
||||
|
||||
Meanwhile [ADR 0037](0037-the-host-applies-it-does-not-decide.md) has given every node a
|
||||
component that does exactly what deploy does — applies state, reads back, reports — and does it
|
||||
continuously rather than once per pipeline. **Two mechanisms now change a node**, and only one
|
||||
of them checks its work.
|
||||
|
||||
## Decision
|
||||
|
||||
**A pipeline ends when the declaration is updated. The node applies it.**
|
||||
|
||||
The three silos become:
|
||||
|
||||
| Silo | Runs | Ends with |
|
||||
|---|---|---|
|
||||
| **build** | once per module | a self-contained artifact |
|
||||
| **publish** | once per module | that artifact addressable — an image by digest, a package in the mesh's repository |
|
||||
| **deploy** | **once, not once per node** | the affected nodes' **declarations updated** to name the new version |
|
||||
|
||||
**Deploy stops sending commands to nodes.** It changes what the control plane says each node
|
||||
should be, which is one write. What happens on the machines is the host's ordinary reconcile,
|
||||
on whatever schedule each node is on.
|
||||
|
||||
### Why this fixes the failure class rather than patching it
|
||||
|
||||
The as-is faults share one shape: **the thing that reported success was not the thing that did
|
||||
the work.** A coordinator dispatching a command can only report on dispatch.
|
||||
|
||||
Under this decision the reporter *is* the applier. The host already refuses to record a resource
|
||||
until it read it back ([ADR 0035](0035-a-picture-is-read-from-what-runs.md)), and already fails
|
||||
the whole apply on one failed step ([ADR 0008](0008-a-failed-step-fails-the-job.md)). A package
|
||||
that 404s cannot go green, because nothing between the package manager and the report has an
|
||||
opportunity to be optimistic.
|
||||
|
||||
**The verify stage disappears as a stage**, which is the strongest evidence for this shape:
|
||||
verification stops being a step that can be omitted from a list, and becomes a property of
|
||||
applying at all.
|
||||
|
||||
### What a pipeline result now means
|
||||
|
||||
The honest answer, and it is different from today's:
|
||||
|
||||
> **The declaration is updated, and here is which nodes have applied it.**
|
||||
|
||||
A pipeline **does not wait for every node**, because a node may be legitimately switched off for
|
||||
a week ([ADR 0036](0036-a-node-is-a-managed-machine.md)) and a delivery mechanism that blocks on
|
||||
a sleeping laptop is one nobody will use. It reports what landed and what has not landed *yet*:
|
||||
|
||||
```
|
||||
delivered declaration updated for 5 nodes
|
||||
applied 3 of 5
|
||||
outstanding 2 — last seen 4 days ago, 20 minutes ago
|
||||
```
|
||||
|
||||
**Outstanding is not failure**, and conflating them is how the old system got a stall with no
|
||||
error anywhere. A node that has not applied yet is a fact with a timestamp, and it resolves
|
||||
itself when the node comes back.
|
||||
|
||||
### The host is delivered the same way as everything else
|
||||
|
||||
The host is tier 0, which makes it tempting to treat as special. It is not:
|
||||
|
||||
1. a push to `mesh-host` builds a binary;
|
||||
2. publish packages it and puts it in **the mesh's own package repository** — which is a
|
||||
directory of files behind the object store and the proxy, so it needs no new machinery;
|
||||
3. deploy updates each node's declaration to name the new version;
|
||||
4. each host applies `package: nox-mesh-host` on its next reconcile, exactly as it applies any
|
||||
other package.
|
||||
|
||||
**No new resource type is needed**, which is the test of whether this is uniform or a special
|
||||
case wearing a uniform. The repository is reachable because a `file` resource put its address in
|
||||
the package manager's configuration — an ordinary declaration, applied by the same host.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **The most consequential boundary in the mesh gets smaller.** The as-is calls the fan-out
|
||||
*"the most consequential boundary in the mesh"* and documents a defect class from the build
|
||||
node having passed through two silos while others had not. **There is no fan-out**: deploy is
|
||||
one write, and the asymmetry it created cannot arise.
|
||||
- **Detection stays the fragile input, and this does not fix it.** *A merge that created no
|
||||
pipeline, and nothing said so* is upstream of everything here and is untouched.
|
||||
- **Rollback becomes a declaration change**, which is a real gain — the previous version is
|
||||
still named in the previous declaration — but nothing here designs how a previous declaration
|
||||
is retained or chosen.
|
||||
- **"Deployed" needs redefining wherever it is used**, because it now means *told*, and the
|
||||
useful fact is *applied on node X at time T*. Anything reporting deployment state has to move
|
||||
to the second, or it will report success for work that has not happened — the exact fault this
|
||||
record is closing, reintroduced at the reporting layer.
|
||||
- **A node offline for a long time applies a large jump at once**, having missed intermediate
|
||||
versions. That is correct — the declaration is a desired state, not a queue of changes — but a
|
||||
machine returning after months applies a very different declaration than it left with, and
|
||||
nothing tests that path.
|
||||
- **The pipeline stops being able to lie and starts being able to be incomplete.** That is a
|
||||
better failure mode and it is still a failure mode: a result that is honest about two
|
||||
outstanding nodes is only useful if somebody looks at it.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0014](0014-build-publish-and-deploy-are-three-silos.md) — the silos this keeps and
|
||||
redefines the third of.
|
||||
- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — the applier that makes this possible.
|
||||
- [ADR 0008](0008-a-failed-step-fails-the-job.md), [ADR 0035](0035-a-picture-is-read-from-what-runs.md) —
|
||||
why the host cannot report success it did not verify.
|
||||
- [`00-as-is/04`](../03-DESIGN/00-as-is/04-delivery.md) — the faults this addresses.
|
||||
@@ -196,24 +196,6 @@ runtime because a declaration changed would stop every container on the node.
|
||||
|
||||
---
|
||||
|
||||
## Updating the host itself
|
||||
|
||||
A declaration too — `package: nox-mesh-host`, at a version. The package manager writes the new
|
||||
binary; the running process is undisturbed, because Unix keeps the running executable's inode.
|
||||
|
||||
**Then the host exits, and the supervisor restarts it on the new binary.** After the apply
|
||||
completes, never during it; only when the executable actually changed; exit zero, so a restart
|
||||
is what happens next rather than a failure a supervisor backs off from.
|
||||
|
||||
**The host never asks the service manager to restart it.** That is the host killing itself
|
||||
part-way through an apply. It stops by finishing.
|
||||
|
||||
**A version-skewed fleet is normal**, because each node restarts when its own apply finishes.
|
||||
What a node reports must be the **running** version, not the installed one, or the mesh will
|
||||
believe an upgrade landed before it took effect.
|
||||
|
||||
---
|
||||
|
||||
## enrolled ⇄ disconnected
|
||||
|
||||
Not a failure. Not degraded. A situation
|
||||
@@ -285,26 +267,184 @@ If `/var/lib/mesh-host/state.json` is lost — a reinstall, a replaced disk —
|
||||
**its record of what it owns**, not its ability to work. It re-enrols, receives the declaration
|
||||
again, and re-applies it.
|
||||
|
||||
**What does not come back is removal.** Resources it applied under an older declaration and no
|
||||
longer holds a record of become unowned: the host will not touch them, because it never touches
|
||||
what it did not create. They sit there, unmanaged, indefinitely.
|
||||
**Without help, what does not come back is removal.** Resources applied under an older
|
||||
declaration, whose record is gone, become unowned: the host will not touch them, because it
|
||||
never touches what it did not create. They would sit there, unmanaged, indefinitely.
|
||||
|
||||
The store is therefore the one piece of node state that matters, and *how it is protected* is
|
||||
not designed.
|
||||
**So the mesh keeps a copy of what each node reports it owns**, refreshed on every apply report,
|
||||
and hands it back on a rebuild — see [Protecting the store](#protecting-the-store). The store
|
||||
remains locally authoritative for *operating*; the copy exists only for this.
|
||||
|
||||
---
|
||||
|
||||
## Open
|
||||
## Upgrading the host
|
||||
|
||||
- **Re-enrolling as the same node.** A machine that lost its identity gets a new token — but
|
||||
whether the mesh treats it as the same node or a new one is an operator's decision today, and
|
||||
nothing supports either.
|
||||
- **Protecting the store.** Above. Its loss is silent and permanent.
|
||||
- **How long disconnected, and who is told.** The fact is not held anywhere.
|
||||
- **Whether a `failed` adoption line still lets adoption complete.** Reopened by
|
||||
[research 012](../../01-RESEARCH/012-the-minimum-viable-node/00-overview.md) and not decided:
|
||||
*flags inform* was decided about conflicts, and a failure is different in kind.
|
||||
- **What a briefing looks like.** It is the first thing a session on a new node sees, which
|
||||
makes it an interface rather than a log.
|
||||
- **Where the enrolment token comes from, operationally.** A person carries it. Nothing says how
|
||||
it is generated, shown, or transported, and it is now the only secret in adoption.
|
||||
The host is delivered like anything else
|
||||
([ADR 0058](../../02-DECISIONS/0058-delivery-ends-in-a-declaration.md)), and this is worth
|
||||
walking through because tier 0 looks like it should be special and is not.
|
||||
|
||||
```
|
||||
push to mesh-host
|
||||
│
|
||||
├─ build go build → one static binary
|
||||
├─ publish packaged, into the mesh's own package repository
|
||||
└─ deploy each node's declaration now names the new version
|
||||
│
|
||||
└─ every host applies it on its next reconcile
|
||||
```
|
||||
|
||||
**Compared with today.** The current pipeline's third silo runs *once per node* and sends each
|
||||
one a command to install and start. That is where the as-is records a package install that
|
||||
404ed while the job went green. Here deploy is **one write** — the declaration changes — and the
|
||||
installing is the host's ordinary work, which reads back before it records anything.
|
||||
|
||||
**The repository is reachable because a declaration made it so.** A `file` resource writes the
|
||||
package manager's configuration pointing at the mesh's repository; a `package` resource names
|
||||
the version. Both ordinary shapes, applied by the same host. **No new resource type**, which is
|
||||
the test of whether this is really uniform.
|
||||
|
||||
### The restart
|
||||
|
||||
```
|
||||
1 pacman installs the new binary the running process is untouched —
|
||||
Unix keeps the running executable's inode
|
||||
2 the host verifies the new binary runs `nox-mesh-host version`, as a subprocess
|
||||
3 it finishes the apply and reports never mid-way
|
||||
4 it exits 0 having finished, not having been stopped
|
||||
5 systemd restarts it on the new binary
|
||||
6 the new host reconciles on start trigger 1, confirming the machine still matches
|
||||
```
|
||||
|
||||
**Step 2 is the one to insist on.** A package can install a binary that does not execute here —
|
||||
wrong architecture, a libc that is not present. Running it once before committing to a restart
|
||||
turns "the node never came back" into "the apply failed and said why". It is the same read-back
|
||||
rule the rest of the host already follows, applied to the one resource that is the host.
|
||||
|
||||
**The host never asks the service manager to restart it.** That is the host stopping itself
|
||||
part-way through an apply. It stops by finishing.
|
||||
|
||||
**A fleet upgrades over an interval, not at an instant**, because each node restarts when its
|
||||
own apply completes. A node must therefore report the version it is **running**, not the one
|
||||
installed — otherwise the mesh believes an upgrade landed at step 1.
|
||||
|
||||
**What is not solved: a new version that crashes on start.** systemd will restart it, back off,
|
||||
and the node is stuck on a binary that does not run. The previous package is still in the
|
||||
package manager's cache, so a person at the machine can downgrade — but there is no automatic
|
||||
rollback, and designing one means the host judging its own health, which is the kind of
|
||||
self-reference the rest of this document avoids. **Named, not solved.**
|
||||
|
||||
---
|
||||
|
||||
## Resolved
|
||||
|
||||
The items this document opened, with the reasoning, because each was open for a reason.
|
||||
|
||||
### Re-enrolling as the same node
|
||||
|
||||
**A token is issued *for* a node record, and that is where the question is answered.**
|
||||
|
||||
```
|
||||
mesh-control token issue --node workstation # this machine is that node again
|
||||
mesh-control token issue --new # a machine the mesh has not seen
|
||||
```
|
||||
|
||||
The host does not need to know which it is. It presents a token and receives an identity; what
|
||||
that identity is bound to was decided when the token was made.
|
||||
|
||||
**Issuing a re-enrolment token revokes the previous identity for that node**, and that is not
|
||||
housekeeping. Two live identities for one node record is the stolen-laptop case with the thief's
|
||||
credentials still valid — the case
|
||||
[`retiring a node`](#enrolled--hosted-retiring-a-node) says is answered by revocation.
|
||||
|
||||
### Protecting the store
|
||||
|
||||
**The host reports what it owns, and the mesh keeps the last report.**
|
||||
|
||||
The store stays locally authoritative — a node operates from its own copy and needs nothing to
|
||||
do so ([ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)). What changes is that
|
||||
the mesh holds a **copy for recovery**, refreshed on every apply report.
|
||||
|
||||
So a node that loses its state file re-enrols, receives both the declaration *and* the record of
|
||||
what it previously owned, and can then remove what is no longer declared. The orphans that used
|
||||
to be permanently stranded are recoverable.
|
||||
|
||||
**This is a backup, never a source.** The host never reads it to decide anything; it is handed
|
||||
back only on a store rebuild, and a node that disagrees with it wins, because the node is the
|
||||
one that can see the machine.
|
||||
|
||||
### How long disconnected, and who is told
|
||||
|
||||
**The mesh records last contact per node; the node records time since it last linked.** Both,
|
||||
because they answer different questions — the mesh's is *have I heard from it*, the node's is
|
||||
*how stale am I*, and a node reporting the second on reconnect is how a long absence gets
|
||||
noticed at all.
|
||||
|
||||
**No threshold and no alarm.** A laptop switched off for three weeks is doing nothing wrong, and
|
||||
a mesh that alerted on it would train people to ignore the alert. It is a **reported fact** —
|
||||
`last seen 4 days ago` beside every node — and what counts as too long is a judgement for
|
||||
whoever is looking, not a constant in the design.
|
||||
|
||||
### Whether a failed adoption line blocks
|
||||
|
||||
**Adoption always completes. A node with a `failed` line is a node, and it is not eligible for
|
||||
assignment until the failure is resolved.**
|
||||
|
||||
This keeps *flags inform, they do not block*
|
||||
([research 012](../../01-RESEARCH/012-the-minimum-viable-node/00-overview.md)) exactly where it
|
||||
was decided — for **conflicts**, where the mesh chose deliberately and the machine works — and
|
||||
gives **failures** the different treatment they need, because a failure is not *we chose* but
|
||||
*we could not*.
|
||||
|
||||
The distinction is between **joining** and **being given work**. Refusing to join makes a
|
||||
machine in use unadoptable, which is the outcome that rule exists to prevent. Placing work on a
|
||||
machine where something the mesh needed never happened produces a module that is installed and
|
||||
does not work — [04-ISSUES/007](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)
|
||||
arriving from the adoption side.
|
||||
|
||||
### What a briefing is
|
||||
|
||||
**A structured document with prose in it**, held in the node's state and reported to the mesh.
|
||||
It is the first thing a session on a new node reads, which makes it an interface.
|
||||
|
||||
```
|
||||
outcome kept derived from the worst line below, never stated separately
|
||||
node workstation
|
||||
adopted 2026-08-27T14:02Z
|
||||
|
||||
ok container runtime docker 27.0, adopted; original config kept at <path>
|
||||
kept storage driver machine has overlay2, the mesh wanted btrfs — machine wins
|
||||
unknown firewall ruleset could not be parsed
|
||||
failed package database locked by another process
|
||||
|
||||
what to look at
|
||||
The storage driver disagreement is preference, not requirement, so nothing is broken.
|
||||
The package database was locked; nothing was installed. Re-run adoption when it is free.
|
||||
```
|
||||
|
||||
**The outcome is computed from the lines**, so a briefing cannot read *fine* while carrying a
|
||||
failed line. Two independently written fields drift, and that drift is the fault this repository
|
||||
keeps cataloguing.
|
||||
|
||||
### Where the enrolment token comes from
|
||||
|
||||
**`mesh-control token issue` prints it once**, to the person running it. Single-use, and it
|
||||
expires whether used or not ([ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md)).
|
||||
|
||||
It is carried by hand — read off a screen, pasted into a terminal. That is not a gap in the
|
||||
design, it is the design: its authenticity comes from the channel it travelled, which is what
|
||||
lets a node verify a mesh it has never spoken to
|
||||
([ADR 0051](../../02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md)). A token emailed,
|
||||
committed, or dropped in shared storage has lost the only property that makes it worth carrying.
|
||||
|
||||
**On the first node it comes from the control plane that was raised two commands ago**, which is
|
||||
the same command against a mesh that is one machine old.
|
||||
|
||||
---
|
||||
|
||||
## Still open
|
||||
|
||||
- **Automatic rollback of a bad host version.** Above. A node can be left on a binary that will
|
||||
not start, and recovering needs a person at the machine.
|
||||
- **How a previous declaration is retained and chosen**, which is what rollback of anything else
|
||||
would use ([ADR 0058](../../02-DECISIONS/0058-delivery-ends-in-a-declaration.md)).
|
||||
- **A node returning after months** applies a very large jump in one go. Correct, and untested.
|
||||
|
||||
Reference in New Issue
Block a user