Files
hq/03-DESIGN/01-to-be/09-the-node-lifecycle.md
T
jschoubben 33a00d5656 Adopt the glossary's vocabulary in the mutable design docs
"control plane" -> controller and "substrate" -> foundation throughout
03-DESIGN, 00-META and the README, with 06-the-control-plane.md and
07-the-substrate.md renamed to 06-the-controller.md and 07-the-foundation.md.
The immutable 02-DECISIONS records keep their original wording (and links to
them are unchanged) — a term retired here may still appear there, which the
glossary explains how to read.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:48:52 +02:00

697 lines
33 KiB
Markdown

---
layer: to-be
status: in-progress
code:
- mesh-host internal/link/run.go
- mesh-host internal/link/enrol.go
- mesh-host packaging/nox-mesh-host-resume.service
- mesh-host packaging/nox-mesh-host-network.sh
- mesh-control internal/token
- mesh-control internal/inventory/nodes.go
updated: 2026-08-31
decisions:
- 02-DECISIONS/0004-a-node-and-how-it-joins.md
- 02-DECISIONS/0005-the-node-host.md
- 02-DECISIONS/0004-a-node-and-how-it-joins.md
- 02-DECISIONS/0004-a-node-and-how-it-joins.md
- 02-DECISIONS/0005-the-node-host.md
- 02-DECISIONS/0004-a-node-and-how-it-joins.md
- 02-DECISIONS/0005-the-node-host.md
- 02-DECISIONS/0010-delivery.md
- 02-DECISIONS/0005-the-node-host.md
- 02-DECISIONS/0005-the-node-host.md
- 02-DECISIONS/0005-the-node-host.md
---
# The node lifecycle
How a Linux machine becomes a node, stays one, and stops being one.
[`05-the-node-host.md`](05-the-node-host.md) describes the host as a component. This describes
it as something that runs for years on a machine somebody else also uses — which is where the
questions that were not being asked live.
## The states
```
unmanaged ──install──► hosted ──enrol──► enrolled ⇄ disconnected
▲ │
└─────release──────┘
```
| State | Has | Can |
|---|---|---|
| **unmanaged** | nothing of ours | — it is a Linux machine |
| **hosted** | the host, no identity | apply a local file, apply its bundle |
| **enrolled** | identity, link, store | everything; this is *a node* |
| **disconnected** | identity, store, no link | hold its machine in the last state it was told |
**Only `enrolled` and `disconnected` are nodes**, and they are the same node in two situations
rather than two kinds of thing
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). **`hosted` is not a
node** — it is a machine with a program on it that has not been told which mesh it belongs to.
There is no state for *the first node*. That is the point of
[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md): the first node walks the
same path, in an unusual order.
---
## unmanaged → hosted: installing
In the machine's own idiom, because the package manager and the init file are the system's
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)):
```
# Alpine — the intended first node
apk add nox-mesh-host
rc-update add nox-mesh-host && rc-service nox-mesh-host start
# Arch
pacman -S nox-mesh-host
systemctl enable --now nox-mesh-host
```
Two lines each, and the init file behind them is four
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)) — it says
*run the launcher at boot* and nothing else, so a third system is transcription rather than a
port.
Or, where there is no repository to install from:
```
curl -fsSL https://<release>/mesh-host-<system>-<version>-x86_64.tar.gz | tar -xz -C /usr/local/bin
```
**The binary is per system as well as per architecture**, because two of its appliers are.
**The tarball must never acquire a dependency**, because the mesh's own package repository is
hosted on the mesh. Any route that needs the mesh in order to install the thing that joins the
mesh is a circle — unusable on a first node, and unusable by whoever is repairing a mesh that is
down, which is exactly when it is wanted.
### The unit it installs
```ini
[Unit]
Description=Novox Mesh node host
After=network-online.target
Wants=network-online.target
[Service]
ExecStart=/usr/lib/nox-mesh-host/launch
Restart=always
RestartSec=5s
StateDirectory=mesh-host
[Install]
WantedBy=multi-user.target
```
**Two lines of policy, and that is deliberate**
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)). The init is
asked to *start this at boot* and *start it again if it exits*, and nothing else. Both are
expressible in OpenRC, runit, s6 and an Android `init.rc`, so porting this file is transcription
rather than design.
**`Restart=always` and not `on-failure`**: the host restarts onto a new binary by exiting
*cleanly*, so a supervisor that only restarts on failure would leave every upgraded node stopped,
having successfully upgraded.
**What the init does not do is decide when to give up.** Counting failed starts and rolling back
lives in the launcher, where it can be tested — `OnFailure=` in a unit file can only be read and
hoped for, and it is the one thing that has to work on a machine where nothing else does.
**The package owns this file. The host never does.** It manages `service` resources, and its own
unit is a service — the temptation is obvious and it ends with a host stopping itself half way
through an apply, leaving a machine with nothing running to fix it. A declaration naming the
host's own unit is **refused**, and that refusal is a test rather than a convention.
The line to hold: **the installation owns the host; the host owns everything else.**
At this point the host is running and **doing nothing**. It has no identity, so there is nobody
to link to and nothing to apply. It answers `profile`, `inventory` and `version`, and waits.
---
## hosted → enrolled: the ordinary case
```
nox-mesh-host enrol --token <one-time token>
```
The token carries **four** things and is carried by a person
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)): the broker's
address, the fingerprint to expect, **the controller's signing identity**, and the right to
join once.
**The fourth is the one this document listed three of.** A node connects to the broker and takes
instruction from the controller behind it, and those are two different identities. Pinning only
the broker would make the controller's authority *transitive* — a compromised broker could then
forge declarations, which, since the host applies whatever the link delivers, is the whole machine.
So the transport is verified once at connect, and **each declaration is verified by its signature,
every time**.
What happens, in order:
1. the host dials the broker at the address in the token, **over the underlay**;
2. it checks the broker's certificate against the pinned fingerprint — *before* sending anything;
3. it presents the one-time secret **and its own public key**, which the mesh records;
4. it reports its `profile` and `inventory` upward;
5. the controller decides what this machine should be, and sends a declaration;
6. the host applies it, reads back, and reports.
**Step 4 is the one that is easy to miss and is what makes step 5 possible.** The controller
cannot decide what a machine should run without knowing what it *can* run — a graphical session,
a container runtime, an architecture. The profile is not a diagnostic; it is the input.
**The node computes nothing about the mesh.** It needs one peer to reach; the whole overlay is
derived centrally and pushed down
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md),
[`08-connectivity.md`](08-connectivity.md)).
### The first declaration is the overlay, and nothing else
**The mesh makes a node reachable before it makes it useful.** Step 5 is not one declaration
carrying everything the node will ever run. It is two, in order:
```
first the overlay — this node's address, its keys, its peers, its names
then everything else — packages, containers, services, files
```
Three reasons, and the third is the one that matters when something goes wrong:
- **It is forced.** A node cannot join the overlay before contacting the mesh, because its
address and peer set are *assigned* — it generates a keypair, publishes the public half, and
receives the rest ([`08-connectivity.md`](08-connectivity.md)). So the overlay is the first
thing the mesh can give it, and it should be.
- **It is what [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) already
says:** *a joining node does the minimum to be reachable, and nothing else.*
- **It is the way back in.** Once the overlay is up, the node is reachable over it — by SSH, by
anything. If a later declaration breaks the machine, there is a route to it that does not
depend on the mesh's control path working. **Sending a large first declaration risks a node
that is broken and unreachable at the same time**, and those two failures are much worse
together than separately.
### Reachable is not the same as having a control surface
Worth stating plainly, because the two rules read as a contradiction and are not.
| | |
|---|---|
| **every node reaches every other node** | over the overlay — SSH, services, ordinary traffic. This is the point of having one |
| **every node consumes from the broker** | its own queue, over its own outbound connection ([ADR 0002](../../02-DECISIONS/0002-nodes-communicate-over-a-broker.md)) |
| **nothing dials a node to control it** | the host has no inbound control surface ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) |
**[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) is about the control
channel, not about network reachability.** What it forbids is a listening thing that accepts
instructions and changes the machine. A node being reachable on the overlay — the whole purpose
of the overlay — is untouched by it, and so is a person opening a shell on it.
The distinction is *who can tell this machine what to be*: only the controller, only over the
link the node opened, only in declarations of known shape.
---
## hosted → enrolled: the first node
The same path, with the mesh built in the middle of it.
```
# 1 — raise the foundation and the controller from the carried bundle
nox-mesh-host reconcile
# 2 — the controller now exists, and issues the first token
mesh-control token issue
# 3 — the machine joins the mesh it just raised
nox-mesh-host enrol --token <token>
```
Step 1 is the bootstrap from [`07-the-foundation.md`](07-the-foundation.md): a container runtime,
then PostgreSQL, then the database, then the schema, then the controller. It needs no identity
because nothing is being asked of anyone — the host is applying a declaration it already
carries, to the machine it is already on.
**After step 3 the first node is not special in any way**, which is the property `adopt.sh` and
the bootstrap script never had. Its specialness lasted two commands.
**And enrolment is exercised on node one.** The path every other node depends on is walked
immediately, against a controller on the same machine, rather than being written and first
used months later on node two.
---
## Two kinds of host
Everything above assumes a machine with an init that runs the host at boot. Not every machine
has one ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)).
| | **resident** | **episodic** |
|---|---|---|
| examples | Alpine, Arch | Android |
| started by | an init, at boot | whatever the platform allows |
| supervised by | the launcher | nothing — the platform decides when it runs |
| the link | held open | opened while it runs |
| being stopped | shutdown, or a failure | **ordinary** |
| shapes | all six | `file`, `directory`, `action` |
| can be the first node | yes | **no** |
**An episodic host being killed is disconnection, not failure.** That is
[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) doing the work it was written
for: reachability is state, not class. Everything the design already does for a laptop that
closes — an authoritative local store, reconcile on start, *last heard from* reported without an
alarm — is what an episodic host needs, at a shorter period.
**It cannot be the first node**, and that is not a limitation to work around. Every step of
raising a foundation is a `package`, a `container` or an `action` against one, and a partial host
refuses the first two. So `mesh-host-android bundle` returns a file that says so rather than an
empty placeholder waiting to be filled in.
**Two things this changes for anything reading the mesh.** *Last heard from* is a much weaker
signal on an episodic host — a healthy phone looks like a dead server — so a reader has to know
which kind it is looking at. And a declaration may take a long time to land, which makes
[ADR 0010](../../02-DECISIONS/0010-delivery.md)'s separation of
*outstanding* from *failed* load-bearing rather than tidy.
**Still open:** how an episodic host is started in practice — an APK with a foreground service,
or Termux with its boot addon — and, first, **what an Android node is for.** A device that can
write files and run commands is not a workload host; it is a presence, or somewhere an agent
runs. Building the start mechanism before deciding that would be building it for nobody.
## Adoption: what happens to what is already there
Adoption is not a state. It is what the **first apply** does when it is told to own something a
machine already has ([research 012](../../01-RESEARCH/012-the-minimum-viable-node/00-overview.md)).
A candidate machine is not empty. It has a package manager, probably a container runtime,
configuration somebody chose. [ADR 0005](../../02-DECISIONS/0005-the-node-host.md)
says the host never touches what it did not create — adoption is the deliberate act of taking
ownership of exactly that, so it is a companion to that rule rather than an exception:
> *never, unless adoption made it the host's* — with adoption **explicit, recorded, and visible
> in what the host says it owns.**
Three rules, all earned:
**The original is kept before anything is written.** A one-way door on a working machine is not
an installation. This is a *never* rule, and it earns that from the worst loss in this record —
a tool acting on a path it did not own.
**On conflict, the machine's configuration wins.** Adoption always completes; the conflict is
flagged and reconciled afterwards. A machine in use keeps working exactly as it did.
**Adoption produces a briefing**, not just a result: what it found, what it took over, and what
it could not resolve — with each line marked `ok`, `kept`, `unknown` or `failed`, and the overall
outcome **derived** from the worst line rather than stated alongside it.
---
## enrolled: what running actually looks like
**Changes are pushed, not polled.** A declaration arrives as a message on the link and the host
applies it then. The link is already open and outbound
([ADR 0002](../../02-DECISIONS/0002-nodes-communicate-over-a-broker.md),
[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) — asking it
repeatedly whether anything has changed would be slower to land *and* constant traffic to learn
nothing.
| Trigger | Kind | |
|---|---|---|
| **a declaration arrives** | **pushed** | the ordinary path — this is how a change lands |
| **start** | event | the machine may have changed while nothing was running |
| **reconnect** | event | declarations may have been missed |
| **every ten minutes** | periodic | **drift, and only drift** |
**Why the timer cannot be an event.** Drift is change the *mesh did not make* — somebody edited
a managed file, a distribution upgrade replaced a config, a container was stopped by hand.
Nothing will ever publish a message about it, because whatever did it is not part of the mesh.
Only looking finds it.
So the two periodic things do different jobs and should not be conflated:
| | direction | answers |
|---|---|---|
| **reconcile timer** | local, looks at the machine | *does this machine still match what it was told?* |
| **heartbeat** | upward, reports to the mesh | *is this node still here, and what is it running?* |
**The heartbeat is what makes silence mean something.** A node with nothing to do sends nothing;
without a heartbeat that is indistinguishable from a node that stopped. With one, *last heard
from* is a fact beside every node — which is what
[how long disconnected](#how-long-disconnected-and-who-is-told) reports and what
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) exists
because a stuck node cannot send.
**Rebooting mid-apply is safe by construction.** The store records each resource *after* it
worked ([ADR 0018](../../02-DECISIONS/0018-a-picture-is-read-from-what-runs.md)), so a host that
dies half way through comes back, finds the completed ones already matching, and applies the
rest. The rule that exists to stop the host lying about what it did also makes it crash-safe.
## Updating what the node holds
An ordinary declaration. Someone assigns a module; the controller recomputes what that node
should be and sends it; the host applies the difference and removes what is no longer declared.
**Removal is not symmetric, and the asymmetry is the design:**
| | on being undeclared |
|---|---|
| file, directory | **removed** |
| container | **removed** — the host created it |
| service | **stopped**; the unit file is not the host's to delete |
| package | **left installed** — *forgotten*, not removed |
| action | **forgotten** — it left nothing the host owns |
The host removes what it *made* and leaves what it merely *configured*. Uninstalling a container
runtime because a declaration changed would stop every container on the node.
---
## enrolled ⇄ disconnected
Not a failure. Not degraded. A situation
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)).
A disconnected node **keeps reconciling against its own store**, so it goes on holding its
machine in the last state it was told to hold. A laptop shut for a week comes back and
reconciles; it does not come back and ask what it is.
What it cannot do: receive new declarations, be granted anything new, or have its certificates
renewed — which is the clock on the whole arrangement
([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)).
**How long it has been disconnected is a fact the mesh must hold**, and nothing holds it today.
Without it, a node running last month's assignments looks exactly like one that is current.
---
## Rescue
The host is still a command-line tool, and that is what rescue is:
```
nox-mesh-host owned # what do you think you own?
nox-mesh-host apply repair.json # apply something by hand, locally
nox-mesh-host profile # what can this machine actually do?
```
`apply FILE` accepts actions, because someone who can write that file and run this binary as
root can already do anything it can. The bound in
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) is on what a
**remote** party may push, not on what a person at the machine may do.
This replaces the three hand-run scripts that exist today — first node, joining, rescue — with
one binary that has always been the same binary.
---
## enrolled → hosted: retiring a node
Two cases, and they are genuinely different.
**Graceful.** The controller sends a final declaration that names nothing. The host removes
what it owns by the table above, reports, and drops its identity. The machine keeps the host
installed and is back to `hosted`. Nothing is left behind that anybody has to remember.
**The node is gone.** Stolen, dead, or simply unreachable. The mesh cannot tell it anything, and
by [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) it will go on reconciling
its last declaration **forever**.
That is the honest consequence of making disconnection ordinary, and the answer is not to make
the host expire. It is that **the node holds nothing that outlives revocation**: its identity is
its own, and every grant it holds is a per-node credential at the provider
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md),
[ADR 0008](../../02-DECISIONS/0008-a-context-owns-its-store.md)). Revoking is done at the
database, the broker, the object store — not on the machine.
So a lost node keeps *running* and stops being able to *reach* anything. That is the best
available outcome and it is worth stating plainly rather than implying the mesh can reach out and
switch a machine off, which it cannot and should not be able to.
---
## Losing the store
Worth its own section because the failure is quiet.
If `/var/lib/mesh-host/state.json` is lost — a reinstall, a replaced disk — the host loses
**its record of what it owns**, not its ability to work. It re-enrols, receives the declaration
again, and re-applies it.
**Without help, what does not come back is removal.** Resources applied under an older
declaration, whose record is gone, become unowned: the host will not touch them, because it
never touches what it did not create. They would sit there, unmanaged, indefinitely.
**So the mesh keeps a copy of what each node reports it owns**, refreshed on every apply report,
and hands it back on a rebuild — see [Protecting the store](#protecting-the-store). The store
remains locally authoritative for *operating*; the copy exists only for this.
---
## Upgrading the host
The host is delivered like anything else
([ADR 0010](../../02-DECISIONS/0010-delivery.md)), and this is worth
walking through because tier 0 looks like it should be special and is not.
```
push to mesh-host
│
├─ build go build → one static binary
├─ publish packaged, into the mesh's own package repository
└─ deploy each node's declaration now names the new version
│
└─ pushed to each node; the host applies it on arrival
(a node that is offline gets it on reconnect)
```
**Compared with today.** The current pipeline's third silo runs *once per node* and sends each
one a command to install and start. That is where the as-is records a package install that
404ed while the job went green. Here deploy is **one write** — the declaration changes — and the
installing is the host's ordinary work, which reads back before it records anything.
**The repository is reachable because a declaration made it so.** A `file` resource writes the
package manager's configuration pointing at the mesh's repository; a `package` resource names
the version. Both ordinary shapes, applied by the same host. **No new resource type**, which is
the test of whether this is really uniform.
### The restart
```
1 pacman installs the new binary the running process is untouched —
Unix keeps the running executable's inode
2 the host verifies the new binary runs `nox-mesh-host version`, as a subprocess
3 it finishes the apply and reports never mid-way
4 it exits 0 having finished, not having been stopped
5 the launcher starts it again on the new binary — it supervises the host
rather than exec'ing it (ADR 0005), so this
needs nothing from the init
6 the new host reconciles on start trigger 1, confirming the machine still matches
```
**Step 2 is the one to insist on.** A package can install a binary that does not execute here —
wrong architecture, a libc that is not present. Running it once before committing to a restart
turns "the node never came back" into "the apply failed and said why". It is the same read-back
rule the rest of the host already follows, applied to the one resource that is the host.
**The host never asks the service manager to restart it.** That is the host stopping itself
part-way through an apply. It stops by finishing.
**A fleet upgrades over an interval, not at an instant**, because each node restarts when its
own apply completes. A node must therefore report the version it is **running**, not the one
installed — otherwise the mesh believes an upgrade landed at step 1.
**A version that crashes on start rolls itself back**
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)).
What the init starts is not the host but a **launcher**, and the launcher is where the policy
lives:
```
init ──► nox-mesh-host-launch ──► nox-mesh-host
├─ halted? say so and stop; a person has to look
├─ count this start attempt
├─ too many, not yet rolled back? roll back, then start
├─ too many, already rolled back? halt — the machine is the problem
└─ otherwise start the host
```
It reinstalls the version recorded in `known-good`, which the host wrote the last time it
completed a reconcile — and the host clears the attempt counter at the same moment, for the same
reason.
**The launcher rather than the init's own features**, because this is the one thing that must
work on a machine where nothing else does. A shell script with a counter can be run against a
stub package manager and asserted; `OnFailure=` in a unit file can only be read and hoped for.
It also means the init is asked for nothing but *start* and *restart*, which every init can
do.
**It rolls back once.** If the previous version also fails, the node stops in a failed state
rather than flapping between two binaries. A second failure is a different diagnosis: the
machine is the problem, not the binary.
**Why this matters more than it looks.** A host that will not start cannot link, and a node that
is not linking looks exactly like a machine somebody switched off — which is the one condition
this design has deliberately decided not to alarm on. Without rollback, a bad release reaches
every node, each one goes quiet, and the mesh reports a fleet of sleeping laptops.
---
## Details that are easy to get wrong
Each of these has a wrong answer that looks reasonable, which is why they are written down
rather than left to be worked out.
### Re-enrolling as the same node
**A token is issued *for* a node record**, and that is where a re-enrolment is decided.
```
mesh-control token issue --node workstation # this machine is that node again
mesh-control token issue --new # a machine the mesh has not seen
```
The host does not need to know which it is. It presents a token and receives an identity; what
that identity is bound to was decided when the token was made.
**Issuing a re-enrolment token revokes the previous identity for that node**, and that is not
housekeeping. Two live identities for one node record is the stolen-laptop case with the thief's
credentials still valid — the case
[`retiring a node`](#enrolled--hosted-retiring-a-node) says is answered by revocation.
### Protecting the store
**The host reports what it owns, and the mesh keeps the last report.**
The store stays locally authoritative — a node operates from its own copy and needs nothing to
do so ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). What changes is that
the mesh holds a **copy for recovery**, refreshed on every apply report.
So a node that loses its state file re-enrols, receives both the declaration *and* the record of
what it previously owned, and can then remove what is no longer declared. The orphans that used
to be permanently stranded are recoverable.
**This is a backup, never a source.** The host never reads it to decide anything; it is handed
back only on a store rebuild, and a node that disagrees with it wins, because the node is the
one that can see the machine.
### How long disconnected, and who is told
**The mesh records last contact per node; the node records time since it last linked.** Both,
because they answer different questions — the mesh's is *have I heard from it*, the node's is
*how stale am I*, and a node reporting the second on reconnect is how a long absence gets
noticed at all.
**No threshold and no alarm.** A laptop switched off for three weeks is doing nothing wrong, and
a mesh that alerted on it would train people to ignore the alert. It is a **reported fact** —
`last seen 4 days ago` beside every node — and what counts as too long is a judgement for
whoever is looking, not a constant in the design.
### Whether a failed adoption line blocks
**Adoption always completes. A node with a `failed` line is a node, and it is not eligible for
assignment until the failure is resolved.**
*Flags inform, they do not block* holds for **conflicts** — where the mesh chose deliberately and
the machine still works. A **failure** is different in kind: not *we chose* but *we could not*,
and it gets different treatment for that reason.
The distinction is between **joining** and **being given work**. Refusing to join makes a
machine in use unadoptable, which is the outcome that rule exists to prevent. Placing work on a
machine where something the mesh needed never happened produces a module that is installed and
does not work — [04-ISSUES/007](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)
arriving from the adoption side.
### What a briefing is
**A structured document with prose in it**, held in the node's state and reported to the mesh.
It is the first thing a session on a new node reads, which makes it an interface.
```
outcome kept derived from the worst line below, never stated separately
node workstation
adopted 2026-08-27T14:02Z
ok container runtime docker 27.0, adopted; original config kept at <path>
kept storage driver machine has overlay2, the mesh wanted btrfs — machine wins
unknown firewall ruleset could not be parsed
failed package database locked by another process
what to look at
The storage driver disagreement is preference, not requirement, so nothing is broken.
The package database was locked; nothing was installed. Re-run adoption when it is free.
```
**The outcome is computed from the lines**, so a briefing cannot read *fine* while carrying a
failed line. Two independently written fields drift, and that drift is the fault this repository
keeps cataloguing.
### Where the enrolment token comes from
**`mesh-control token issue` prints it once**, to the person running it. Single-use, and it
expires whether used or not ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)).
It is carried by hand — read off a screen, pasted into a terminal. That is the design rather than
a gap in it: its authenticity comes from the channel it travelled, which is what
lets a node verify a mesh it has never spoken to
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). A token emailed,
committed, or dropped in shared storage has lost the only property that makes it worth carrying.
**On the first node it comes from the controller that was raised two commands ago**, which is
the same command against a mesh that is one machine old.
---
## Still open
- ~~**Automatic rollback of a bad host version.**~~ **Resolved** by
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md): a launcher
counts failed starts and rolls back — shipped by the package, not the host binary, because a
binary that will not start cannot recover itself. It rolls back once; a second failure means
the machine is the problem, not the binary.
- **How a previous declaration is retained and chosen**, which is what rollback of anything else
would use ([ADR 0010](../../02-DECISIONS/0010-delivery.md)).
- **A node returning after months** applies a very large jump in one go. Correct, and untested.
## Going away and coming back
*2026-08-31, from being asked whether a machine that drops off has to be adopted again.*
**It does not, and nothing about it expires.** A disconnected node is the same node in a different
situation ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) — reachability is state,
not class. The node holds its own identity and the mesh holds the public half; there is no lease,
no timeout, and nothing that lapses while a machine is shut. A laptop closed for a week comes back
and reconnects, and a declaration sent while it was away is waiting for it.
**The only thing that forces re-enrolment is a machine losing its own key** — a reinstall, an image
re-cloned. That is deliberate: the mesh then reports what it can no longer seal to rather than
delivering blobs the machine cannot open.
### The gap that was left, and why it mattered
**A machine that suspends does not know it has been disconnected, and neither does anything else.**
After a resume the socket looks perfectly healthy from inside the process: no error, no close,
because nothing has tried to send anything yet. Heartbeats discover it twenty or thirty seconds
later.
For those twenty or thirty seconds the node believes it is in the mesh and is not — and *absence
must never be indistinguishable from a failure to answer* is the rule this whole design is built
on. It recovered on its own, which is why this was a quality gap rather than a fault. It was still
the machine waiting to be told something it already knew.
**So the machine says so.** Waking, and changing network, both rouse the host.
| | |
|---|---|
| **it ends the current attempt** | not the wait after it — the process is not waiting, it is sitting inside a connection that will never return |
| **by signal, not by anything listening** | a socket for this would be a control surface on every machine, in exchange for saving twenty seconds, and the security argument rests on there not being one |
| **two rouses at once are one** | a machine suspending and resuming repeatedly must not build a backlog of reconnections to work through |
| **the backoff is not reset** | being roused says the machine changed, not that whatever refused the connection has stopped. A laptop woken on a network with no route would otherwise retry at full speed for as long as somebody keeps opening the lid |
| **`down` does not rouse** | the link is already gone, reconnecting will fail, and the backoff exists for exactly that |
*Checked by holding a link that never returns on its own — which is precisely what a suspended
connection is — rousing it, and requiring the attempt to end and another to begin. And by running
the dispatcher against every event a network manager emits, requiring it to act on the ones that
change where packets go and on no others.*