diff --git a/02-DECISIONS/0058-delivery-ends-in-a-declaration.md b/02-DECISIONS/0058-delivery-ends-in-a-declaration.md new file mode 100644 index 0000000..2e89c38 --- /dev/null +++ b/02-DECISIONS/0058-delivery-ends-in-a-declaration.md @@ -0,0 +1,130 @@ +--- +status: proposed +date: 2026-08-27 +deciders: jochen +reconstructed: false +extends: 0014-build-publish-and-deploy-are-three-silos.md +--- + +# 58. Delivery ends in a declaration, not in a push to a node + +## Context + +Today a push produces a pipeline with three silos, and the third — **deploy** — runs *once per +module per node*, sending a command to every assigned node telling it to install, configure, +start and verify ([`00-as-is/04`](../03-DESIGN/00-as-is/04-delivery.md)). + +That third silo is where the as-is document records the most damage: + +- *"A green pipeline proves transport, not effect."* The stages report a message was dispatched + and accepted, not that anything is running. +- A service reported started when the container command merely returned. An image pull failure + that did not fail the deploy. **A package install that 404ed from every mirror while the job + went green** ([04-ISSUES/001](../04-ISSUES/001-failed-package-install-reports-success/00-report.md)). +- A node left on old code after a failed download, with a version marker that had already + advanced. +- A verify stage built to close the gap, and *never scheduled*, because the coordinator's stage + list did not include it. + +Meanwhile [ADR 0037](0037-the-host-applies-it-does-not-decide.md) has given every node a +component that does exactly what deploy does — applies state, reads back, reports — and does it +continuously rather than once per pipeline. **Two mechanisms now change a node**, and only one +of them checks its work. + +## Decision + +**A pipeline ends when the declaration is updated. The node applies it.** + +The three silos become: + +| Silo | Runs | Ends with | +|---|---|---| +| **build** | once per module | a self-contained artifact | +| **publish** | once per module | that artifact addressable — an image by digest, a package in the mesh's repository | +| **deploy** | **once, not once per node** | the affected nodes' **declarations updated** to name the new version | + +**Deploy stops sending commands to nodes.** It changes what the control plane says each node +should be, which is one write. What happens on the machines is the host's ordinary reconcile, +on whatever schedule each node is on. + +### Why this fixes the failure class rather than patching it + +The as-is faults share one shape: **the thing that reported success was not the thing that did +the work.** A coordinator dispatching a command can only report on dispatch. + +Under this decision the reporter *is* the applier. The host already refuses to record a resource +until it read it back ([ADR 0035](0035-a-picture-is-read-from-what-runs.md)), and already fails +the whole apply on one failed step ([ADR 0008](0008-a-failed-step-fails-the-job.md)). A package +that 404s cannot go green, because nothing between the package manager and the report has an +opportunity to be optimistic. + +**The verify stage disappears as a stage**, which is the strongest evidence for this shape: +verification stops being a step that can be omitted from a list, and becomes a property of +applying at all. + +### What a pipeline result now means + +The honest answer, and it is different from today's: + +> **The declaration is updated, and here is which nodes have applied it.** + +A pipeline **does not wait for every node**, because a node may be legitimately switched off for +a week ([ADR 0036](0036-a-node-is-a-managed-machine.md)) and a delivery mechanism that blocks on +a sleeping laptop is one nobody will use. It reports what landed and what has not landed *yet*: + +``` +delivered declaration updated for 5 nodes +applied 3 of 5 +outstanding 2 — last seen 4 days ago, 20 minutes ago +``` + +**Outstanding is not failure**, and conflating them is how the old system got a stall with no +error anywhere. A node that has not applied yet is a fact with a timestamp, and it resolves +itself when the node comes back. + +### The host is delivered the same way as everything else + +The host is tier 0, which makes it tempting to treat as special. It is not: + +1. a push to `mesh-host` builds a binary; +2. publish packages it and puts it in **the mesh's own package repository** — which is a + directory of files behind the object store and the proxy, so it needs no new machinery; +3. deploy updates each node's declaration to name the new version; +4. each host applies `package: nox-mesh-host` on its next reconcile, exactly as it applies any + other package. + +**No new resource type is needed**, which is the test of whether this is uniform or a special +case wearing a uniform. The repository is reachable because a `file` resource put its address in +the package manager's configuration — an ordinary declaration, applied by the same host. + +## Consequences + +- **The most consequential boundary in the mesh gets smaller.** The as-is calls the fan-out + *"the most consequential boundary in the mesh"* and documents a defect class from the build + node having passed through two silos while others had not. **There is no fan-out**: deploy is + one write, and the asymmetry it created cannot arise. +- **Detection stays the fragile input, and this does not fix it.** *A merge that created no + pipeline, and nothing said so* is upstream of everything here and is untouched. +- **Rollback becomes a declaration change**, which is a real gain — the previous version is + still named in the previous declaration — but nothing here designs how a previous declaration + is retained or chosen. +- **"Deployed" needs redefining wherever it is used**, because it now means *told*, and the + useful fact is *applied on node X at time T*. Anything reporting deployment state has to move + to the second, or it will report success for work that has not happened — the exact fault this + record is closing, reintroduced at the reporting layer. +- **A node offline for a long time applies a large jump at once**, having missed intermediate + versions. That is correct — the declaration is a desired state, not a queue of changes — but a + machine returning after months applies a very different declaration than it left with, and + nothing tests that path. +- **The pipeline stops being able to lie and starts being able to be incomplete.** That is a + better failure mode and it is still a failure mode: a result that is honest about two + outstanding nodes is only useful if somebody looks at it. + +## References + +- [ADR 0014](0014-build-publish-and-deploy-are-three-silos.md) — the silos this keeps and + redefines the third of. +- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — the applier that makes this possible. +- [ADR 0008](0008-a-failed-step-fails-the-job.md), [ADR 0035](0035-a-picture-is-read-from-what-runs.md) — + why the host cannot report success it did not verify. +- [`00-as-is/04`](../03-DESIGN/00-as-is/04-delivery.md) — the faults this addresses. diff --git a/03-DESIGN/01-to-be/09-the-node-lifecycle.md b/03-DESIGN/01-to-be/09-the-node-lifecycle.md index 6db99f5..1bd3cc7 100644 --- a/03-DESIGN/01-to-be/09-the-node-lifecycle.md +++ b/03-DESIGN/01-to-be/09-the-node-lifecycle.md @@ -196,24 +196,6 @@ runtime because a declaration changed would stop every container on the node. --- -## Updating the host itself - -A declaration too — `package: nox-mesh-host`, at a version. The package manager writes the new -binary; the running process is undisturbed, because Unix keeps the running executable's inode. - -**Then the host exits, and the supervisor restarts it on the new binary.** After the apply -completes, never during it; only when the executable actually changed; exit zero, so a restart -is what happens next rather than a failure a supervisor backs off from. - -**The host never asks the service manager to restart it.** That is the host killing itself -part-way through an apply. It stops by finishing. - -**A version-skewed fleet is normal**, because each node restarts when its own apply finishes. -What a node reports must be the **running** version, not the installed one, or the mesh will -believe an upgrade landed before it took effect. - ---- - ## enrolled ⇄ disconnected Not a failure. Not degraded. A situation @@ -285,26 +267,184 @@ If `/var/lib/mesh-host/state.json` is lost — a reinstall, a replaced disk — **its record of what it owns**, not its ability to work. It re-enrols, receives the declaration again, and re-applies it. -**What does not come back is removal.** Resources it applied under an older declaration and no -longer holds a record of become unowned: the host will not touch them, because it never touches -what it did not create. They sit there, unmanaged, indefinitely. +**Without help, what does not come back is removal.** Resources applied under an older +declaration, whose record is gone, become unowned: the host will not touch them, because it +never touches what it did not create. They would sit there, unmanaged, indefinitely. -The store is therefore the one piece of node state that matters, and *how it is protected* is -not designed. +**So the mesh keeps a copy of what each node reports it owns**, refreshed on every apply report, +and hands it back on a rebuild — see [Protecting the store](#protecting-the-store). The store +remains locally authoritative for *operating*; the copy exists only for this. --- -## Open +## Upgrading the host -- **Re-enrolling as the same node.** A machine that lost its identity gets a new token — but - whether the mesh treats it as the same node or a new one is an operator's decision today, and - nothing supports either. -- **Protecting the store.** Above. Its loss is silent and permanent. -- **How long disconnected, and who is told.** The fact is not held anywhere. -- **Whether a `failed` adoption line still lets adoption complete.** Reopened by - [research 012](../../01-RESEARCH/012-the-minimum-viable-node/00-overview.md) and not decided: - *flags inform* was decided about conflicts, and a failure is different in kind. -- **What a briefing looks like.** It is the first thing a session on a new node sees, which - makes it an interface rather than a log. -- **Where the enrolment token comes from, operationally.** A person carries it. Nothing says how - it is generated, shown, or transported, and it is now the only secret in adoption. +The host is delivered like anything else +([ADR 0058](../../02-DECISIONS/0058-delivery-ends-in-a-declaration.md)), and this is worth +walking through because tier 0 looks like it should be special and is not. + +``` +push to mesh-host + │ + ├─ build go build → one static binary + ├─ publish packaged, into the mesh's own package repository + └─ deploy each node's declaration now names the new version + │ + └─ every host applies it on its next reconcile +``` + +**Compared with today.** The current pipeline's third silo runs *once per node* and sends each +one a command to install and start. That is where the as-is records a package install that +404ed while the job went green. Here deploy is **one write** — the declaration changes — and the +installing is the host's ordinary work, which reads back before it records anything. + +**The repository is reachable because a declaration made it so.** A `file` resource writes the +package manager's configuration pointing at the mesh's repository; a `package` resource names +the version. Both ordinary shapes, applied by the same host. **No new resource type**, which is +the test of whether this is really uniform. + +### The restart + +``` +1 pacman installs the new binary the running process is untouched — + Unix keeps the running executable's inode +2 the host verifies the new binary runs `nox-mesh-host version`, as a subprocess +3 it finishes the apply and reports never mid-way +4 it exits 0 having finished, not having been stopped +5 systemd restarts it on the new binary +6 the new host reconciles on start trigger 1, confirming the machine still matches +``` + +**Step 2 is the one to insist on.** A package can install a binary that does not execute here — +wrong architecture, a libc that is not present. Running it once before committing to a restart +turns "the node never came back" into "the apply failed and said why". It is the same read-back +rule the rest of the host already follows, applied to the one resource that is the host. + +**The host never asks the service manager to restart it.** That is the host stopping itself +part-way through an apply. It stops by finishing. + +**A fleet upgrades over an interval, not at an instant**, because each node restarts when its +own apply completes. A node must therefore report the version it is **running**, not the one +installed — otherwise the mesh believes an upgrade landed at step 1. + +**What is not solved: a new version that crashes on start.** systemd will restart it, back off, +and the node is stuck on a binary that does not run. The previous package is still in the +package manager's cache, so a person at the machine can downgrade — but there is no automatic +rollback, and designing one means the host judging its own health, which is the kind of +self-reference the rest of this document avoids. **Named, not solved.** + +--- + +## Resolved + +The items this document opened, with the reasoning, because each was open for a reason. + +### Re-enrolling as the same node + +**A token is issued *for* a node record, and that is where the question is answered.** + +``` +mesh-control token issue --node workstation # this machine is that node again +mesh-control token issue --new # a machine the mesh has not seen +``` + +The host does not need to know which it is. It presents a token and receives an identity; what +that identity is bound to was decided when the token was made. + +**Issuing a re-enrolment token revokes the previous identity for that node**, and that is not +housekeeping. Two live identities for one node record is the stolen-laptop case with the thief's +credentials still valid — the case +[`retiring a node`](#enrolled--hosted-retiring-a-node) says is answered by revocation. + +### Protecting the store + +**The host reports what it owns, and the mesh keeps the last report.** + +The store stays locally authoritative — a node operates from its own copy and needs nothing to +do so ([ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)). What changes is that +the mesh holds a **copy for recovery**, refreshed on every apply report. + +So a node that loses its state file re-enrols, receives both the declaration *and* the record of +what it previously owned, and can then remove what is no longer declared. The orphans that used +to be permanently stranded are recoverable. + +**This is a backup, never a source.** The host never reads it to decide anything; it is handed +back only on a store rebuild, and a node that disagrees with it wins, because the node is the +one that can see the machine. + +### How long disconnected, and who is told + +**The mesh records last contact per node; the node records time since it last linked.** Both, +because they answer different questions — the mesh's is *have I heard from it*, the node's is +*how stale am I*, and a node reporting the second on reconnect is how a long absence gets +noticed at all. + +**No threshold and no alarm.** A laptop switched off for three weeks is doing nothing wrong, and +a mesh that alerted on it would train people to ignore the alert. It is a **reported fact** — +`last seen 4 days ago` beside every node — and what counts as too long is a judgement for +whoever is looking, not a constant in the design. + +### Whether a failed adoption line blocks + +**Adoption always completes. A node with a `failed` line is a node, and it is not eligible for +assignment until the failure is resolved.** + +This keeps *flags inform, they do not block* +([research 012](../../01-RESEARCH/012-the-minimum-viable-node/00-overview.md)) exactly where it +was decided — for **conflicts**, where the mesh chose deliberately and the machine works — and +gives **failures** the different treatment they need, because a failure is not *we chose* but +*we could not*. + +The distinction is between **joining** and **being given work**. Refusing to join makes a +machine in use unadoptable, which is the outcome that rule exists to prevent. Placing work on a +machine where something the mesh needed never happened produces a module that is installed and +does not work — [04-ISSUES/007](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md) +arriving from the adoption side. + +### What a briefing is + +**A structured document with prose in it**, held in the node's state and reported to the mesh. +It is the first thing a session on a new node reads, which makes it an interface. + +``` +outcome kept derived from the worst line below, never stated separately +node workstation +adopted 2026-08-27T14:02Z + + ok container runtime docker 27.0, adopted; original config kept at + kept storage driver machine has overlay2, the mesh wanted btrfs — machine wins + unknown firewall ruleset could not be parsed + failed package database locked by another process + +what to look at + The storage driver disagreement is preference, not requirement, so nothing is broken. + The package database was locked; nothing was installed. Re-run adoption when it is free. +``` + +**The outcome is computed from the lines**, so a briefing cannot read *fine* while carrying a +failed line. Two independently written fields drift, and that drift is the fault this repository +keeps cataloguing. + +### Where the enrolment token comes from + +**`mesh-control token issue` prints it once**, to the person running it. Single-use, and it +expires whether used or not ([ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md)). + +It is carried by hand — read off a screen, pasted into a terminal. That is not a gap in the +design, it is the design: its authenticity comes from the channel it travelled, which is what +lets a node verify a mesh it has never spoken to +([ADR 0051](../../02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md)). A token emailed, +committed, or dropped in shared storage has lost the only property that makes it worth carrying. + +**On the first node it comes from the control plane that was raised two commands ago**, which is +the same command against a mesh that is one machine old. + +--- + +## Still open + +- **Automatic rollback of a bad host version.** Above. A node can be left on a binary that will + not start, and recovering needs a person at the machine. +- **How a previous declaration is retained and chosen**, which is what rollback of anything else + would use ([ADR 0058](../../02-DECISIONS/0058-delivery-ends-in-a-declaration.md)). +- **A node returning after months** applies a very large jump in one go. Correct, and untested.