Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24
@@ -0,0 +1,130 @@
|
||||
---
|
||||
status: proposed
|
||||
date: 2026-08-27
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 0014-build-publish-and-deploy-are-three-silos.md
|
||||
---
|
||||
|
||||
# 58. Delivery ends in a declaration, not in a push to a node
|
||||
|
||||
## Context
|
||||
|
||||
Today a push produces a pipeline with three silos, and the third — **deploy** — runs *once per
|
||||
module per node*, sending a command to every assigned node telling it to install, configure,
|
||||
start and verify ([`00-as-is/04`](../03-DESIGN/00-as-is/04-delivery.md)).
|
||||
|
||||
That third silo is where the as-is document records the most damage:
|
||||
|
||||
- *"A green pipeline proves transport, not effect."* The stages report a message was dispatched
|
||||
and accepted, not that anything is running.
|
||||
- A service reported started when the container command merely returned. An image pull failure
|
||||
that did not fail the deploy. **A package install that 404ed from every mirror while the job
|
||||
went green** ([04-ISSUES/001](../04-ISSUES/001-failed-package-install-reports-success/00-report.md)).
|
||||
- A node left on old code after a failed download, with a version marker that had already
|
||||
advanced.
|
||||
- A verify stage built to close the gap, and *never scheduled*, because the coordinator's stage
|
||||
list did not include it.
|
||||
|
||||
Meanwhile [ADR 0037](0037-the-host-applies-it-does-not-decide.md) has given every node a
|
||||
component that does exactly what deploy does — applies state, reads back, reports — and does it
|
||||
continuously rather than once per pipeline. **Two mechanisms now change a node**, and only one
|
||||
of them checks its work.
|
||||
|
||||
## Decision
|
||||
|
||||
**A pipeline ends when the declaration is updated. The node applies it.**
|
||||
|
||||
The three silos become:
|
||||
|
||||
| Silo | Runs | Ends with |
|
||||
|---|---|---|
|
||||
| **build** | once per module | a self-contained artifact |
|
||||
| **publish** | once per module | that artifact addressable — an image by digest, a package in the mesh's repository |
|
||||
| **deploy** | **once, not once per node** | the affected nodes' **declarations updated** to name the new version |
|
||||
|
||||
**Deploy stops sending commands to nodes.** It changes what the control plane says each node
|
||||
should be, which is one write. What happens on the machines is the host's ordinary reconcile,
|
||||
on whatever schedule each node is on.
|
||||
|
||||
### Why this fixes the failure class rather than patching it
|
||||
|
||||
The as-is faults share one shape: **the thing that reported success was not the thing that did
|
||||
the work.** A coordinator dispatching a command can only report on dispatch.
|
||||
|
||||
Under this decision the reporter *is* the applier. The host already refuses to record a resource
|
||||
until it read it back ([ADR 0035](0035-a-picture-is-read-from-what-runs.md)), and already fails
|
||||
the whole apply on one failed step ([ADR 0008](0008-a-failed-step-fails-the-job.md)). A package
|
||||
that 404s cannot go green, because nothing between the package manager and the report has an
|
||||
opportunity to be optimistic.
|
||||
|
||||
**The verify stage disappears as a stage**, which is the strongest evidence for this shape:
|
||||
verification stops being a step that can be omitted from a list, and becomes a property of
|
||||
applying at all.
|
||||
|
||||
### What a pipeline result now means
|
||||
|
||||
The honest answer, and it is different from today's:
|
||||
|
||||
> **The declaration is updated, and here is which nodes have applied it.**
|
||||
|
||||
A pipeline **does not wait for every node**, because a node may be legitimately switched off for
|
||||
a week ([ADR 0036](0036-a-node-is-a-managed-machine.md)) and a delivery mechanism that blocks on
|
||||
a sleeping laptop is one nobody will use. It reports what landed and what has not landed *yet*:
|
||||
|
||||
```
|
||||
delivered declaration updated for 5 nodes
|
||||
applied 3 of 5
|
||||
outstanding 2 — last seen 4 days ago, 20 minutes ago
|
||||
```
|
||||
|
||||
**Outstanding is not failure**, and conflating them is how the old system got a stall with no
|
||||
error anywhere. A node that has not applied yet is a fact with a timestamp, and it resolves
|
||||
itself when the node comes back.
|
||||
|
||||
### The host is delivered the same way as everything else
|
||||
|
||||
The host is tier 0, which makes it tempting to treat as special. It is not:
|
||||
|
||||
1. a push to `mesh-host` builds a binary;
|
||||
2. publish packages it and puts it in **the mesh's own package repository** — which is a
|
||||
directory of files behind the object store and the proxy, so it needs no new machinery;
|
||||
3. deploy updates each node's declaration to name the new version;
|
||||
4. each host applies `package: nox-mesh-host` on its next reconcile, exactly as it applies any
|
||||
other package.
|
||||
|
||||
**No new resource type is needed**, which is the test of whether this is uniform or a special
|
||||
case wearing a uniform. The repository is reachable because a `file` resource put its address in
|
||||
the package manager's configuration — an ordinary declaration, applied by the same host.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **The most consequential boundary in the mesh gets smaller.** The as-is calls the fan-out
|
||||
*"the most consequential boundary in the mesh"* and documents a defect class from the build
|
||||
node having passed through two silos while others had not. **There is no fan-out**: deploy is
|
||||
one write, and the asymmetry it created cannot arise.
|
||||
- **Detection stays the fragile input, and this does not fix it.** *A merge that created no
|
||||
pipeline, and nothing said so* is upstream of everything here and is untouched.
|
||||
- **Rollback becomes a declaration change**, which is a real gain — the previous version is
|
||||
still named in the previous declaration — but nothing here designs how a previous declaration
|
||||
is retained or chosen.
|
||||
- **"Deployed" needs redefining wherever it is used**, because it now means *told*, and the
|
||||
useful fact is *applied on node X at time T*. Anything reporting deployment state has to move
|
||||
to the second, or it will report success for work that has not happened — the exact fault this
|
||||
record is closing, reintroduced at the reporting layer.
|
||||
- **A node offline for a long time applies a large jump at once**, having missed intermediate
|
||||
versions. That is correct — the declaration is a desired state, not a queue of changes — but a
|
||||
machine returning after months applies a very different declaration than it left with, and
|
||||
nothing tests that path.
|
||||
- **The pipeline stops being able to lie and starts being able to be incomplete.** That is a
|
||||
better failure mode and it is still a failure mode: a result that is honest about two
|
||||
outstanding nodes is only useful if somebody looks at it.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0014](0014-build-publish-and-deploy-are-three-silos.md) — the silos this keeps and
|
||||
redefines the third of.
|
||||
- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — the applier that makes this possible.
|
||||
- [ADR 0008](0008-a-failed-step-fails-the-job.md), [ADR 0035](0035-a-picture-is-read-from-what-runs.md) —
|
||||
why the host cannot report success it did not verify.
|
||||
- [`00-as-is/04`](../03-DESIGN/00-as-is/04-delivery.md) — the faults this addresses.
|
||||
@@ -196,24 +196,6 @@ runtime because a declaration changed would stop every container on the node.
|
||||
|
||||
---
|
||||
|
||||
## Updating the host itself
|
||||
|
||||
A declaration too — `package: nox-mesh-host`, at a version. The package manager writes the new
|
||||
binary; the running process is undisturbed, because Unix keeps the running executable's inode.
|
||||
|
||||
**Then the host exits, and the supervisor restarts it on the new binary.** After the apply
|
||||
completes, never during it; only when the executable actually changed; exit zero, so a restart
|
||||
is what happens next rather than a failure a supervisor backs off from.
|
||||
|
||||
**The host never asks the service manager to restart it.** That is the host killing itself
|
||||
part-way through an apply. It stops by finishing.
|
||||
|
||||
**A version-skewed fleet is normal**, because each node restarts when its own apply finishes.
|
||||
What a node reports must be the **running** version, not the installed one, or the mesh will
|
||||
believe an upgrade landed before it took effect.
|
||||
|
||||
---
|
||||
|
||||
## enrolled ⇄ disconnected
|
||||
|
||||
Not a failure. Not degraded. A situation
|
||||
@@ -285,26 +267,184 @@ If `/var/lib/mesh-host/state.json` is lost — a reinstall, a replaced disk —
|
||||
**its record of what it owns**, not its ability to work. It re-enrols, receives the declaration
|
||||
again, and re-applies it.
|
||||
|
||||
**What does not come back is removal.** Resources it applied under an older declaration and no
|
||||
longer holds a record of become unowned: the host will not touch them, because it never touches
|
||||
what it did not create. They sit there, unmanaged, indefinitely.
|
||||
**Without help, what does not come back is removal.** Resources applied under an older
|
||||
declaration, whose record is gone, become unowned: the host will not touch them, because it
|
||||
never touches what it did not create. They would sit there, unmanaged, indefinitely.
|
||||
|
||||
The store is therefore the one piece of node state that matters, and *how it is protected* is
|
||||
not designed.
|
||||
**So the mesh keeps a copy of what each node reports it owns**, refreshed on every apply report,
|
||||
and hands it back on a rebuild — see [Protecting the store](#protecting-the-store). The store
|
||||
remains locally authoritative for *operating*; the copy exists only for this.
|
||||
|
||||
---
|
||||
|
||||
## Open
|
||||
## Upgrading the host
|
||||
|
||||
- **Re-enrolling as the same node.** A machine that lost its identity gets a new token — but
|
||||
whether the mesh treats it as the same node or a new one is an operator's decision today, and
|
||||
nothing supports either.
|
||||
- **Protecting the store.** Above. Its loss is silent and permanent.
|
||||
- **How long disconnected, and who is told.** The fact is not held anywhere.
|
||||
- **Whether a `failed` adoption line still lets adoption complete.** Reopened by
|
||||
[research 012](../../01-RESEARCH/012-the-minimum-viable-node/00-overview.md) and not decided:
|
||||
*flags inform* was decided about conflicts, and a failure is different in kind.
|
||||
- **What a briefing looks like.** It is the first thing a session on a new node sees, which
|
||||
makes it an interface rather than a log.
|
||||
- **Where the enrolment token comes from, operationally.** A person carries it. Nothing says how
|
||||
it is generated, shown, or transported, and it is now the only secret in adoption.
|
||||
The host is delivered like anything else
|
||||
([ADR 0058](../../02-DECISIONS/0058-delivery-ends-in-a-declaration.md)), and this is worth
|
||||
walking through because tier 0 looks like it should be special and is not.
|
||||
|
||||
```
|
||||
push to mesh-host
|
||||
│
|
||||
├─ build go build → one static binary
|
||||
├─ publish packaged, into the mesh's own package repository
|
||||
└─ deploy each node's declaration now names the new version
|
||||
│
|
||||
└─ every host applies it on its next reconcile
|
||||
```
|
||||
|
||||
**Compared with today.** The current pipeline's third silo runs *once per node* and sends each
|
||||
one a command to install and start. That is where the as-is records a package install that
|
||||
404ed while the job went green. Here deploy is **one write** — the declaration changes — and the
|
||||
installing is the host's ordinary work, which reads back before it records anything.
|
||||
|
||||
**The repository is reachable because a declaration made it so.** A `file` resource writes the
|
||||
package manager's configuration pointing at the mesh's repository; a `package` resource names
|
||||
the version. Both ordinary shapes, applied by the same host. **No new resource type**, which is
|
||||
the test of whether this is really uniform.
|
||||
|
||||
### The restart
|
||||
|
||||
```
|
||||
1 pacman installs the new binary the running process is untouched —
|
||||
Unix keeps the running executable's inode
|
||||
2 the host verifies the new binary runs `nox-mesh-host version`, as a subprocess
|
||||
3 it finishes the apply and reports never mid-way
|
||||
4 it exits 0 having finished, not having been stopped
|
||||
5 systemd restarts it on the new binary
|
||||
6 the new host reconciles on start trigger 1, confirming the machine still matches
|
||||
```
|
||||
|
||||
**Step 2 is the one to insist on.** A package can install a binary that does not execute here —
|
||||
wrong architecture, a libc that is not present. Running it once before committing to a restart
|
||||
turns "the node never came back" into "the apply failed and said why". It is the same read-back
|
||||
rule the rest of the host already follows, applied to the one resource that is the host.
|
||||
|
||||
**The host never asks the service manager to restart it.** That is the host stopping itself
|
||||
part-way through an apply. It stops by finishing.
|
||||
|
||||
**A fleet upgrades over an interval, not at an instant**, because each node restarts when its
|
||||
own apply completes. A node must therefore report the version it is **running**, not the one
|
||||
installed — otherwise the mesh believes an upgrade landed at step 1.
|
||||
|
||||
**What is not solved: a new version that crashes on start.** systemd will restart it, back off,
|
||||
and the node is stuck on a binary that does not run. The previous package is still in the
|
||||
package manager's cache, so a person at the machine can downgrade — but there is no automatic
|
||||
rollback, and designing one means the host judging its own health, which is the kind of
|
||||
self-reference the rest of this document avoids. **Named, not solved.**
|
||||
|
||||
---
|
||||
|
||||
## Resolved
|
||||
|
||||
The items this document opened, with the reasoning, because each was open for a reason.
|
||||
|
||||
### Re-enrolling as the same node
|
||||
|
||||
**A token is issued *for* a node record, and that is where the question is answered.**
|
||||
|
||||
```
|
||||
mesh-control token issue --node workstation # this machine is that node again
|
||||
mesh-control token issue --new # a machine the mesh has not seen
|
||||
```
|
||||
|
||||
The host does not need to know which it is. It presents a token and receives an identity; what
|
||||
that identity is bound to was decided when the token was made.
|
||||
|
||||
**Issuing a re-enrolment token revokes the previous identity for that node**, and that is not
|
||||
housekeeping. Two live identities for one node record is the stolen-laptop case with the thief's
|
||||
credentials still valid — the case
|
||||
[`retiring a node`](#enrolled--hosted-retiring-a-node) says is answered by revocation.
|
||||
|
||||
### Protecting the store
|
||||
|
||||
**The host reports what it owns, and the mesh keeps the last report.**
|
||||
|
||||
The store stays locally authoritative — a node operates from its own copy and needs nothing to
|
||||
do so ([ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)). What changes is that
|
||||
the mesh holds a **copy for recovery**, refreshed on every apply report.
|
||||
|
||||
So a node that loses its state file re-enrols, receives both the declaration *and* the record of
|
||||
what it previously owned, and can then remove what is no longer declared. The orphans that used
|
||||
to be permanently stranded are recoverable.
|
||||
|
||||
**This is a backup, never a source.** The host never reads it to decide anything; it is handed
|
||||
back only on a store rebuild, and a node that disagrees with it wins, because the node is the
|
||||
one that can see the machine.
|
||||
|
||||
### How long disconnected, and who is told
|
||||
|
||||
**The mesh records last contact per node; the node records time since it last linked.** Both,
|
||||
because they answer different questions — the mesh's is *have I heard from it*, the node's is
|
||||
*how stale am I*, and a node reporting the second on reconnect is how a long absence gets
|
||||
noticed at all.
|
||||
|
||||
**No threshold and no alarm.** A laptop switched off for three weeks is doing nothing wrong, and
|
||||
a mesh that alerted on it would train people to ignore the alert. It is a **reported fact** —
|
||||
`last seen 4 days ago` beside every node — and what counts as too long is a judgement for
|
||||
whoever is looking, not a constant in the design.
|
||||
|
||||
### Whether a failed adoption line blocks
|
||||
|
||||
**Adoption always completes. A node with a `failed` line is a node, and it is not eligible for
|
||||
assignment until the failure is resolved.**
|
||||
|
||||
This keeps *flags inform, they do not block*
|
||||
([research 012](../../01-RESEARCH/012-the-minimum-viable-node/00-overview.md)) exactly where it
|
||||
was decided — for **conflicts**, where the mesh chose deliberately and the machine works — and
|
||||
gives **failures** the different treatment they need, because a failure is not *we chose* but
|
||||
*we could not*.
|
||||
|
||||
The distinction is between **joining** and **being given work**. Refusing to join makes a
|
||||
machine in use unadoptable, which is the outcome that rule exists to prevent. Placing work on a
|
||||
machine where something the mesh needed never happened produces a module that is installed and
|
||||
does not work — [04-ISSUES/007](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)
|
||||
arriving from the adoption side.
|
||||
|
||||
### What a briefing is
|
||||
|
||||
**A structured document with prose in it**, held in the node's state and reported to the mesh.
|
||||
It is the first thing a session on a new node reads, which makes it an interface.
|
||||
|
||||
```
|
||||
outcome kept derived from the worst line below, never stated separately
|
||||
node workstation
|
||||
adopted 2026-08-27T14:02Z
|
||||
|
||||
ok container runtime docker 27.0, adopted; original config kept at <path>
|
||||
kept storage driver machine has overlay2, the mesh wanted btrfs — machine wins
|
||||
unknown firewall ruleset could not be parsed
|
||||
failed package database locked by another process
|
||||
|
||||
what to look at
|
||||
The storage driver disagreement is preference, not requirement, so nothing is broken.
|
||||
The package database was locked; nothing was installed. Re-run adoption when it is free.
|
||||
```
|
||||
|
||||
**The outcome is computed from the lines**, so a briefing cannot read *fine* while carrying a
|
||||
failed line. Two independently written fields drift, and that drift is the fault this repository
|
||||
keeps cataloguing.
|
||||
|
||||
### Where the enrolment token comes from
|
||||
|
||||
**`mesh-control token issue` prints it once**, to the person running it. Single-use, and it
|
||||
expires whether used or not ([ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md)).
|
||||
|
||||
It is carried by hand — read off a screen, pasted into a terminal. That is not a gap in the
|
||||
design, it is the design: its authenticity comes from the channel it travelled, which is what
|
||||
lets a node verify a mesh it has never spoken to
|
||||
([ADR 0051](../../02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md)). A token emailed,
|
||||
committed, or dropped in shared storage has lost the only property that makes it worth carrying.
|
||||
|
||||
**On the first node it comes from the control plane that was raised two commands ago**, which is
|
||||
the same command against a mesh that is one machine old.
|
||||
|
||||
---
|
||||
|
||||
## Still open
|
||||
|
||||
- **Automatic rollback of a bad host version.** Above. A node can be left on a binary that will
|
||||
not start, and recovering needs a person at the machine.
|
||||
- **How a previous declaration is retained and chosen**, which is what rollback of anything else
|
||||
would use ([ADR 0058](../../02-DECISIONS/0058-delivery-ends-in-a-declaration.md)).
|
||||
- **A node returning after months** applies a very large jump in one go. Correct, and untested.
|
||||
|
||||
Reference in New Issue
Block a user