Two corrections and one new decision, all from Jochen catching things. Pushed, not polled. I described updates as landing "on the next reconcile", which reads as polling and is not the design. A declaration arrives as a message on a link that is already open; the host applies it then. Polling over an existing connection would be slower to land AND constant traffic to learn nothing. The timer is for drift and nothing else, and it cannot be replaced by an event for a definitional reason: drift is change the mesh did not make -- somebody edited a managed file, a distribution upgrade replaced a config -- so nothing will ever publish a message about it. Only looking finds it. Separated the heartbeat from the reconcile timer, which I had been conflating. They point in opposite directions and answer different questions: the timer looks at the machine and asks whether it still matches; the heartbeat reports upward and is what makes silence mean something. A node with nothing to do sends nothing, and without a heartbeat that is indistinguishable from a node that stopped. 0059 -- a host that cannot start is rolled back by the service manager. I had left this open on the grounds that recovery meant the host judging its own health. That objection does not survive being asked properly: a keepalive is something else judging the host. The watchdog must be local, because nothing dials a node and a host that cannot start cannot report -- so it is the service manager, which is already there. The failure it prevents is sharper than "the node is down": a host that will not start looks exactly like a machine somebody switched off, which is the one condition this design has deliberately decided not to alarm on. So a bad release reaches every node, each goes quiet, and the mesh reports a fleet of sleeping laptops. Confirmed means started and completed one reconcile -- deliberately not "the link is up", or a laptop on a train would roll itself back. The rollback is a script shipped by the package, not a host subcommand, because a binary that will not start cannot be its own recovery. It rolls back once: a second failure means the machine is the problem, not the binary. Also refined the records checker, which produced a false positive: a proposed record may extend another proposed one, because decisions are drafted in chains and the alternative is marking things accepted to satisfy a check. An accepted document resting on a proposed record still fails, and that was verified. 0057, 0058 and 0059 are all proposed.
485 lines
22 KiB
Markdown
485 lines
22 KiB
Markdown
---
|
|
layer: to-be
|
|
status: designed
|
|
code: []
|
|
updated: 2026-08-27
|
|
decisions:
|
|
- 02-DECISIONS/0036-a-node-is-a-managed-machine.md
|
|
- 02-DECISIONS/0037-the-host-applies-it-does-not-decide.md
|
|
- 02-DECISIONS/0038-a-node-joins-by-linking-first.md
|
|
- 02-DECISIONS/0039-the-link-is-the-security-boundary.md
|
|
- 02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md
|
|
- 02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md
|
|
---
|
|
|
|
# The node lifecycle
|
|
|
|
How a Linux machine becomes a node, stays one, and stops being one.
|
|
|
|
[`05-the-node-host.md`](05-the-node-host.md) describes the host as a component. This describes
|
|
it as something that runs for years on a machine somebody else also uses — which is where the
|
|
questions that were not being asked live.
|
|
|
|
## The states
|
|
|
|
```
|
|
unmanaged ──install──► hosted ──enrol──► enrolled ⇄ disconnected
|
|
▲ │
|
|
└─────release──────┘
|
|
```
|
|
|
|
| State | Has | Can |
|
|
|---|---|---|
|
|
| **unmanaged** | nothing of ours | — it is a Linux machine |
|
|
| **hosted** | the host, no identity | apply a local file, apply its bundle |
|
|
| **enrolled** | identity, link, store | everything; this is *a node* |
|
|
| **disconnected** | identity, store, no link | hold its machine in the last state it was told |
|
|
|
|
**Only `enrolled` and `disconnected` are nodes**, and they are the same node in two situations
|
|
rather than two kinds of thing
|
|
([ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)). **`hosted` is not a
|
|
node** — it is a machine with a program on it that has not been told which mesh it belongs to.
|
|
|
|
There is no state for *the first node*. That is the point of
|
|
[ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md): the first node walks the
|
|
same path, in an unusual order.
|
|
|
|
---
|
|
|
|
## unmanaged → hosted: installing
|
|
|
|
```
|
|
pacman -S nox-mesh-host
|
|
systemctl enable --now nox-mesh-host
|
|
```
|
|
|
|
Or, where there is no repository to install from:
|
|
|
|
```
|
|
curl -fsSL https://<release>/mesh-host-<version>-x86_64.tar.gz | tar -xz -C /usr/local/bin
|
|
```
|
|
|
|
**The tarball must never acquire a dependency**, because the mesh's own package repository is
|
|
hosted on the mesh. Any route that needs the mesh in order to install the thing that joins the
|
|
mesh is a circle — unusable on a first node, and unusable by whoever is repairing a mesh that is
|
|
down, which is exactly when it is wanted.
|
|
|
|
At this point the host is running and **doing nothing**. It has no identity, so there is nobody
|
|
to link to and nothing to apply. It answers `profile`, `inventory` and `version`, and waits.
|
|
|
|
---
|
|
|
|
## hosted → enrolled: the ordinary case
|
|
|
|
```
|
|
nox-mesh-host enrol --token <one-time token>
|
|
```
|
|
|
|
The token carries three things and is carried by a person
|
|
([ADR 0051](../../02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md)): the broker's
|
|
address, the fingerprint to expect, and the right to join once.
|
|
|
|
What happens, in order:
|
|
|
|
1. the host dials the broker at the address in the token, **over the underlay**;
|
|
2. it checks the broker's certificate against the pinned fingerprint — *before* sending anything;
|
|
3. it presents the one-time secret and receives its **own durable identity**;
|
|
4. it reports its `profile` and `inventory` upward;
|
|
5. the control plane decides what this machine should be, and sends a declaration;
|
|
6. the host applies it, reads back, and reports.
|
|
|
|
**Step 4 is the one that is easy to miss and is what makes step 5 possible.** The control plane
|
|
cannot decide what a machine should run without knowing what it *can* run — a graphical session,
|
|
a container runtime, an architecture. The profile is not a diagnostic; it is the input.
|
|
|
|
**The node computes nothing about the mesh.** It needs one peer to reach; the whole overlay is
|
|
derived centrally and pushed down
|
|
([ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md),
|
|
[`08-connectivity.md`](08-connectivity.md)).
|
|
|
|
---
|
|
|
|
## hosted → enrolled: the first node
|
|
|
|
The same path, with the mesh built in the middle of it.
|
|
|
|
```
|
|
# 1 — raise the substrate and the control plane from the carried bundle
|
|
nox-mesh-host reconcile
|
|
|
|
# 2 — the control plane now exists, and issues the first token
|
|
mesh-control token issue
|
|
|
|
# 3 — the machine joins the mesh it just raised
|
|
nox-mesh-host enrol --token <token>
|
|
```
|
|
|
|
Step 1 is the bootstrap from [`07-the-substrate.md`](07-the-substrate.md): Docker, then
|
|
PostgreSQL, then the database, then the schema, then the control plane. It needs no identity
|
|
because nothing is being asked of anyone — the host is applying a declaration it already
|
|
carries, to the machine it is already on.
|
|
|
|
**After step 3 the first node is not special in any way**, which is the property `adopt.sh` and
|
|
the bootstrap script never had. Its specialness lasted two commands.
|
|
|
|
**And enrolment is exercised on node one.** The path every other node depends on is walked
|
|
immediately, against a control plane on the same machine, rather than being written and first
|
|
used months later on node two.
|
|
|
|
---
|
|
|
|
## Adoption: what happens to what is already there
|
|
|
|
Adoption is not a state. It is what the **first apply** does when it is told to own something a
|
|
machine already has ([research 012](../../01-RESEARCH/012-the-minimum-viable-node/00-overview.md)).
|
|
|
|
A candidate machine is not empty. It has a package manager, probably a container runtime,
|
|
configuration somebody chose. [ADR 0043](../../02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md)
|
|
says the host never touches what it did not create — adoption is the deliberate act of taking
|
|
ownership of exactly that, so it is a companion to that rule rather than an exception:
|
|
|
|
> *never, unless adoption made it the host's* — with adoption **explicit, recorded, and visible
|
|
> in what the host says it owns.**
|
|
|
|
Three rules, all earned:
|
|
|
|
**The original is kept before anything is written.** A one-way door on a working machine is not
|
|
an installation. This is a *never* rule, and it earns that from the worst loss in this record —
|
|
a tool acting on a path it did not own.
|
|
|
|
**On conflict, the machine's configuration wins.** Adoption always completes; the conflict is
|
|
flagged and reconciled afterwards. A machine in use keeps working exactly as it did.
|
|
|
|
**Adoption produces a briefing**, not just a result: what it found, what it took over, and what
|
|
it could not resolve — with each line marked `ok`, `kept`, `unknown` or `failed`, and the overall
|
|
outcome **derived** from the worst line rather than stated alongside it.
|
|
|
|
---
|
|
|
|
## enrolled: what running actually looks like
|
|
|
|
**Changes are pushed, not polled.** A declaration arrives as a message on the link and the host
|
|
applies it then. The link is already open and outbound
|
|
([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md),
|
|
[ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md)) — asking it
|
|
repeatedly whether anything has changed would be slower to land *and* constant traffic to learn
|
|
nothing.
|
|
|
|
| Trigger | Kind | |
|
|
|---|---|---|
|
|
| **a declaration arrives** | **pushed** | the ordinary path — this is how a change lands |
|
|
| **start** | event | the machine may have changed while nothing was running |
|
|
| **reconnect** | event | declarations may have been missed |
|
|
| **every ten minutes** | periodic | **drift, and only drift** |
|
|
|
|
**Why the timer cannot be an event.** Drift is change the *mesh did not make* — somebody edited
|
|
a managed file, a distribution upgrade replaced a config, a container was stopped by hand.
|
|
Nothing will ever publish a message about it, because whatever did it is not part of the mesh.
|
|
Only looking finds it.
|
|
|
|
So the two periodic things do different jobs and should not be conflated:
|
|
|
|
| | direction | answers |
|
|
|---|---|---|
|
|
| **reconcile timer** | local, looks at the machine | *does this machine still match what it was told?* |
|
|
| **heartbeat** | upward, reports to the mesh | *is this node still here, and what is it running?* |
|
|
|
|
**The heartbeat is what makes silence mean something.** A node with nothing to do sends nothing;
|
|
without a heartbeat that is indistinguishable from a node that stopped. With one, *last heard
|
|
from* is a fact beside every node — which is what
|
|
[how long disconnected](#how-long-disconnected-and-who-is-told) reports and what
|
|
[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md) exists because
|
|
a stuck node cannot send.
|
|
|
|
**Rebooting mid-apply is safe by construction.** The store records each resource *after* it
|
|
worked ([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)), so a host that
|
|
dies half way through comes back, finds the completed ones already matching, and applies the
|
|
rest. The rule that exists to stop the host lying about what it did also makes it crash-safe.
|
|
|
|
## Updating what the node holds
|
|
|
|
An ordinary declaration. Someone assigns a module; the control plane recomputes what that node
|
|
should be and sends it; the host applies the difference and removes what is no longer declared.
|
|
|
|
**Removal is not symmetric, and the asymmetry is the design:**
|
|
|
|
| | on being undeclared |
|
|
|---|---|
|
|
| file, directory | **removed** |
|
|
| container | **removed** — the host created it |
|
|
| service | **stopped**; the unit file is not the host's to delete |
|
|
| package | **left installed** — *forgotten*, not removed |
|
|
| action | **forgotten** — it left nothing the host owns |
|
|
|
|
The host removes what it *made* and leaves what it merely *configured*. Uninstalling a container
|
|
runtime because a declaration changed would stop every container on the node.
|
|
|
|
---
|
|
|
|
## enrolled ⇄ disconnected
|
|
|
|
Not a failure. Not degraded. A situation
|
|
([ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)).
|
|
|
|
A disconnected node **keeps reconciling against its own store**, so it goes on holding its
|
|
machine in the last state it was told to hold. A laptop shut for a week comes back and
|
|
reconciles; it does not come back and ask what it is.
|
|
|
|
What it cannot do: receive new declarations, be granted anything new, or have its certificates
|
|
renewed — which is the clock on the whole arrangement
|
|
([ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md)).
|
|
|
|
**How long it has been disconnected is a fact the mesh must hold**, and nothing holds it today.
|
|
Without it, a node running last month's assignments looks exactly like one that is current.
|
|
|
|
---
|
|
|
|
## Rescue
|
|
|
|
The host is still a command-line tool, and that is what rescue is:
|
|
|
|
```
|
|
nox-mesh-host owned # what do you think you own?
|
|
nox-mesh-host apply repair.json # apply something by hand, locally
|
|
nox-mesh-host profile # what can this machine actually do?
|
|
```
|
|
|
|
`apply FILE` accepts actions, because someone who can write that file and run this binary as
|
|
root can already do anything it can. The bound in
|
|
[ADR 0047](../../02-DECISIONS/0047-the-bundle-may-carry-actions-the-link-may-not.md) is on what a
|
|
**remote** party may push, not on what a person at the machine may do.
|
|
|
|
This replaces the three hand-run scripts that exist today — first node, joining, rescue — with
|
|
one binary that has always been the same binary.
|
|
|
|
---
|
|
|
|
## enrolled → hosted: retiring a node
|
|
|
|
Two cases, and they are genuinely different.
|
|
|
|
**Graceful.** The control plane sends a final declaration that names nothing. The host removes
|
|
what it owns by the table above, reports, and drops its identity. The machine keeps the host
|
|
installed and is back to `hosted`. Nothing is left behind that anybody has to remember.
|
|
|
|
**The node is gone.** Stolen, dead, or simply unreachable. The mesh cannot tell it anything, and
|
|
by [ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md) it will go on reconciling
|
|
its last declaration **forever**.
|
|
|
|
That is the honest consequence of making disconnection ordinary, and the answer is not to make
|
|
the host expire. It is that **the node holds nothing that outlives revocation**: its identity is
|
|
its own, and every grant it holds is a per-node credential at the provider
|
|
([ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md),
|
|
[ADR 0045](../../02-DECISIONS/0045-a-context-owns-its-store.md)). Revoking is done at the
|
|
database, the broker, the object store — not on the machine.
|
|
|
|
So a lost node keeps *running* and stops being able to *reach* anything. That is the best
|
|
available outcome and it is worth stating plainly rather than implying the mesh can reach out and
|
|
switch a machine off, which it cannot and should not be able to.
|
|
|
|
---
|
|
|
|
## Losing the store
|
|
|
|
Worth its own section because the failure is quiet.
|
|
|
|
If `/var/lib/mesh-host/state.json` is lost — a reinstall, a replaced disk — the host loses
|
|
**its record of what it owns**, not its ability to work. It re-enrols, receives the declaration
|
|
again, and re-applies it.
|
|
|
|
**Without help, what does not come back is removal.** Resources applied under an older
|
|
declaration, whose record is gone, become unowned: the host will not touch them, because it
|
|
never touches what it did not create. They would sit there, unmanaged, indefinitely.
|
|
|
|
**So the mesh keeps a copy of what each node reports it owns**, refreshed on every apply report,
|
|
and hands it back on a rebuild — see [Protecting the store](#protecting-the-store). The store
|
|
remains locally authoritative for *operating*; the copy exists only for this.
|
|
|
|
---
|
|
|
|
## Upgrading the host
|
|
|
|
The host is delivered like anything else
|
|
([ADR 0058](../../02-DECISIONS/0058-delivery-ends-in-a-declaration.md)), and this is worth
|
|
walking through because tier 0 looks like it should be special and is not.
|
|
|
|
```
|
|
push to mesh-host
|
|
│
|
|
├─ build go build → one static binary
|
|
├─ publish packaged, into the mesh's own package repository
|
|
└─ deploy each node's declaration now names the new version
|
|
│
|
|
└─ pushed to each node; the host applies it on arrival
|
|
(a node that is offline gets it on reconnect)
|
|
```
|
|
|
|
**Compared with today.** The current pipeline's third silo runs *once per node* and sends each
|
|
one a command to install and start. That is where the as-is records a package install that
|
|
404ed while the job went green. Here deploy is **one write** — the declaration changes — and the
|
|
installing is the host's ordinary work, which reads back before it records anything.
|
|
|
|
**The repository is reachable because a declaration made it so.** A `file` resource writes the
|
|
package manager's configuration pointing at the mesh's repository; a `package` resource names
|
|
the version. Both ordinary shapes, applied by the same host. **No new resource type**, which is
|
|
the test of whether this is really uniform.
|
|
|
|
### The restart
|
|
|
|
```
|
|
1 pacman installs the new binary the running process is untouched —
|
|
Unix keeps the running executable's inode
|
|
2 the host verifies the new binary runs `nox-mesh-host version`, as a subprocess
|
|
3 it finishes the apply and reports never mid-way
|
|
4 it exits 0 having finished, not having been stopped
|
|
5 systemd restarts it on the new binary
|
|
6 the new host reconciles on start trigger 1, confirming the machine still matches
|
|
```
|
|
|
|
**Step 2 is the one to insist on.** A package can install a binary that does not execute here —
|
|
wrong architecture, a libc that is not present. Running it once before committing to a restart
|
|
turns "the node never came back" into "the apply failed and said why". It is the same read-back
|
|
rule the rest of the host already follows, applied to the one resource that is the host.
|
|
|
|
**The host never asks the service manager to restart it.** That is the host stopping itself
|
|
part-way through an apply. It stops by finishing.
|
|
|
|
**A fleet upgrades over an interval, not at an instant**, because each node restarts when its
|
|
own apply completes. A node must therefore report the version it is **running**, not the one
|
|
installed — otherwise the mesh believes an upgrade landed at step 1.
|
|
|
|
**A version that crashes on start rolls itself back**
|
|
([ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md)). The service
|
|
manager gives up after three failures in two minutes and runs a rollback script — shipped by the
|
|
package rather than being a host subcommand, because a binary that will not start cannot be its
|
|
own recovery. It reinstalls the version recorded in `known-good`, which the host wrote the last
|
|
time it completed a reconcile.
|
|
|
|
**It rolls back once.** If the previous version also fails, the node stops in a failed state
|
|
rather than flapping between two binaries. A second failure is a different diagnosis: the
|
|
machine is the problem, not the binary.
|
|
|
|
**Why this matters more than it looks.** A host that will not start cannot link, and a node that
|
|
is not linking looks exactly like a machine somebody switched off — which is the one condition
|
|
this design has deliberately decided not to alarm on. Without rollback, a bad release reaches
|
|
every node, each one goes quiet, and the mesh reports a fleet of sleeping laptops.
|
|
|
|
---
|
|
|
|
## Resolved
|
|
|
|
The items this document opened, with the reasoning, because each was open for a reason.
|
|
|
|
### Re-enrolling as the same node
|
|
|
|
**A token is issued *for* a node record, and that is where the question is answered.**
|
|
|
|
```
|
|
mesh-control token issue --node workstation # this machine is that node again
|
|
mesh-control token issue --new # a machine the mesh has not seen
|
|
```
|
|
|
|
The host does not need to know which it is. It presents a token and receives an identity; what
|
|
that identity is bound to was decided when the token was made.
|
|
|
|
**Issuing a re-enrolment token revokes the previous identity for that node**, and that is not
|
|
housekeeping. Two live identities for one node record is the stolen-laptop case with the thief's
|
|
credentials still valid — the case
|
|
[`retiring a node`](#enrolled--hosted-retiring-a-node) says is answered by revocation.
|
|
|
|
### Protecting the store
|
|
|
|
**The host reports what it owns, and the mesh keeps the last report.**
|
|
|
|
The store stays locally authoritative — a node operates from its own copy and needs nothing to
|
|
do so ([ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)). What changes is that
|
|
the mesh holds a **copy for recovery**, refreshed on every apply report.
|
|
|
|
So a node that loses its state file re-enrols, receives both the declaration *and* the record of
|
|
what it previously owned, and can then remove what is no longer declared. The orphans that used
|
|
to be permanently stranded are recoverable.
|
|
|
|
**This is a backup, never a source.** The host never reads it to decide anything; it is handed
|
|
back only on a store rebuild, and a node that disagrees with it wins, because the node is the
|
|
one that can see the machine.
|
|
|
|
### How long disconnected, and who is told
|
|
|
|
**The mesh records last contact per node; the node records time since it last linked.** Both,
|
|
because they answer different questions — the mesh's is *have I heard from it*, the node's is
|
|
*how stale am I*, and a node reporting the second on reconnect is how a long absence gets
|
|
noticed at all.
|
|
|
|
**No threshold and no alarm.** A laptop switched off for three weeks is doing nothing wrong, and
|
|
a mesh that alerted on it would train people to ignore the alert. It is a **reported fact** —
|
|
`last seen 4 days ago` beside every node — and what counts as too long is a judgement for
|
|
whoever is looking, not a constant in the design.
|
|
|
|
### Whether a failed adoption line blocks
|
|
|
|
**Adoption always completes. A node with a `failed` line is a node, and it is not eligible for
|
|
assignment until the failure is resolved.**
|
|
|
|
This keeps *flags inform, they do not block*
|
|
([research 012](../../01-RESEARCH/012-the-minimum-viable-node/00-overview.md)) exactly where it
|
|
was decided — for **conflicts**, where the mesh chose deliberately and the machine works — and
|
|
gives **failures** the different treatment they need, because a failure is not *we chose* but
|
|
*we could not*.
|
|
|
|
The distinction is between **joining** and **being given work**. Refusing to join makes a
|
|
machine in use unadoptable, which is the outcome that rule exists to prevent. Placing work on a
|
|
machine where something the mesh needed never happened produces a module that is installed and
|
|
does not work — [04-ISSUES/007](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)
|
|
arriving from the adoption side.
|
|
|
|
### What a briefing is
|
|
|
|
**A structured document with prose in it**, held in the node's state and reported to the mesh.
|
|
It is the first thing a session on a new node reads, which makes it an interface.
|
|
|
|
```
|
|
outcome kept derived from the worst line below, never stated separately
|
|
node workstation
|
|
adopted 2026-08-27T14:02Z
|
|
|
|
ok container runtime docker 27.0, adopted; original config kept at <path>
|
|
kept storage driver machine has overlay2, the mesh wanted btrfs — machine wins
|
|
unknown firewall ruleset could not be parsed
|
|
failed package database locked by another process
|
|
|
|
what to look at
|
|
The storage driver disagreement is preference, not requirement, so nothing is broken.
|
|
The package database was locked; nothing was installed. Re-run adoption when it is free.
|
|
```
|
|
|
|
**The outcome is computed from the lines**, so a briefing cannot read *fine* while carrying a
|
|
failed line. Two independently written fields drift, and that drift is the fault this repository
|
|
keeps cataloguing.
|
|
|
|
### Where the enrolment token comes from
|
|
|
|
**`mesh-control token issue` prints it once**, to the person running it. Single-use, and it
|
|
expires whether used or not ([ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md)).
|
|
|
|
It is carried by hand — read off a screen, pasted into a terminal. That is not a gap in the
|
|
design, it is the design: its authenticity comes from the channel it travelled, which is what
|
|
lets a node verify a mesh it has never spoken to
|
|
([ADR 0051](../../02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md)). A token emailed,
|
|
committed, or dropped in shared storage has lost the only property that makes it worth carrying.
|
|
|
|
**On the first node it comes from the control plane that was raised two commands ago**, which is
|
|
the same command against a mesh that is one machine old.
|
|
|
|
---
|
|
|
|
## Still open
|
|
|
|
- ~~**Automatic rollback of a bad host version.**~~ **Resolved** by
|
|
[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md): the service
|
|
manager gives up after three failures and runs a rollback script — shipped by the package, not
|
|
the host binary, because a binary that will not start cannot recover itself. It rolls back
|
|
once; a second failure means the machine is the problem, not the binary.
|
|
- **How a previous declaration is retained and chosen**, which is what rollback of anything else
|
|
would use ([ADR 0058](../../02-DECISIONS/0058-delivery-ends-in-a-declaration.md)).
|
|
- **A node returning after months** applies a very large jump in one go. Correct, and untested.
|