Design the node lifecycle end to end
The host was described as a component and never as something that runs for years on a machine somebody else also uses. 09 covers every state a machine can be in and every transition between them. Four states: unmanaged, hosted, enrolled, disconnected. Only the last two are nodes, and they are the same node in two situations. `hosted` -- the host installed but never told which mesh it belongs to -- had no name before and is where a machine sits between the two adoption commands. Things that were unclear and now are not: The first node walks the same path in an unusual order: reconcile from the bundle, the control plane it just raised issues a token, enrol against it. Its specialness lasts two commands. A side effect worth having -- enrolment is exercised on node one, rather than being written and first used on node two. Enrolment reports profile and inventory BEFORE the control plane decides anything. The profile is the input to that decision, not a diagnostic; the control plane cannot decide what a machine should run without knowing what it can run. Rebooting mid-apply is safe by construction. The store records each resource after it worked, so a host that dies half way through comes back and applies the rest. The rule that stops the host lying about what it did also makes it crash-safe. Retiring splits in two. Graceful is a final empty declaration. A node that is gone will reconcile its last declaration forever -- the honest consequence of making disconnection ordinary. The answer is not to make the host expire but that the node holds nothing that outlives revocation: every grant is a per-node credential revoked at the provider. A lost node keeps running and stops being able to reach anything. Said plainly rather than implying the mesh can switch a machine off, which it cannot and should not. Losing the store is quiet and permanent, so it gets its own section. The host re-enrols and re-applies fine; what does not come back is removal, because resources it no longer has a record of become unowned and sit there indefinitely. Also corrects 0057, which said the mesh must not upgrade the host at all. That conflated two acts. Replacing the binary is safe -- Unix keeps the running inode. Stopping the unit is not. So the host may apply a package naming itself, and restarts by finishing its apply and exiting cleanly, letting the supervisor start it on the new binary. It never asks the service manager to restart it. That makes a fleet-wide host upgrade an ordinary declaration, which the first draft gave up on. 0057 remains proposed.
This commit is contained in:
@@ -72,14 +72,33 @@ and neither can a node whose mesh is down — which is exactly when somebody is
|
||||
**Any path that requires the mesh to install the thing that joins the mesh is a circle**, so the
|
||||
tarball is the path that is never allowed to acquire a dependency.
|
||||
|
||||
### The mesh does not upgrade the host
|
||||
### The host may replace its own binary; it may not stop its own unit
|
||||
|
||||
A running process replacing its own binary and restarting mid-apply is the self-management
|
||||
problem wearing a different hat. **Upgrading the host is an act on the machine**, by the package
|
||||
manager, not a declaration the host applies to itself.
|
||||
The first draft of this record said the mesh must not upgrade the host at all. That was too
|
||||
broad, and it conflated two different acts.
|
||||
|
||||
Recorded as a limit rather than a plan: it means a fleet-wide host upgrade is not currently a
|
||||
mesh operation, and that will be felt.
|
||||
**Replacing the binary is safe.** Unix keeps the running executable's inode open, so a package
|
||||
upgrade writes a new file and the running process continues on the old one, undisturbed.
|
||||
**Stopping the unit is what is unsafe** — that is the host killing itself part-way through an
|
||||
apply, leaving a machine with nothing running to finish or fix it.
|
||||
|
||||
So the host may apply a `package` naming itself. What it must never do is ask the service
|
||||
manager to restart it.
|
||||
|
||||
**The restart happens by exiting, not by asking.** When the host notices its own executable has
|
||||
been replaced, it finishes the apply it is in, reports what it did, and **exits cleanly**. The
|
||||
supervisor's `Restart=always` starts it again, on the new binary. Nothing stops the host; the
|
||||
host stops, having finished.
|
||||
|
||||
Three conditions, and they are the whole safety argument:
|
||||
|
||||
- **after** the apply completes and its outcomes are recorded — never mid-way;
|
||||
- **only** when the executable actually changed, which Linux reports plainly: a replaced
|
||||
`/proc/self/exe` reads as the old path marked deleted;
|
||||
- **exit zero**, so a restart is what a supervisor does next rather than a failure it backs off
|
||||
from.
|
||||
|
||||
This makes a fleet-wide host upgrade an ordinary declaration, which the first draft gave up.
|
||||
|
||||
### It reconciles on four triggers
|
||||
|
||||
@@ -115,8 +134,14 @@ host already knows it owns.
|
||||
- **The timer makes drift visible and also makes it loud.** A resource the host cannot apply
|
||||
will now fail every ten minutes rather than once. That is correct and it needs somewhere to go
|
||||
other than a log nobody reads — which is `observability`'s, and it does not exist yet.
|
||||
- **Host upgrades are outside the mesh**, so a fleet-wide upgrade is currently manual. Nothing
|
||||
here solves it and it should not be solved by giving the host a self-upgrade path.
|
||||
- **A host upgrade is an ordinary declaration**, which is worth the care it needs: the exit
|
||||
path is the only place the host deliberately stops, and a bug there is a node that restarts in
|
||||
a loop or never comes back. It wants a test that the host does **not** exit when its binary is
|
||||
unchanged, as much as one that it does when it changed.
|
||||
- **A version-skewed fleet is now normal and needs saying.** Nodes restart onto the new binary
|
||||
at whatever moment their apply finishes, so "the fleet is upgraded" is a range rather than an
|
||||
instant. What a node reports as its version must be the **running** one, not the installed
|
||||
one, or the mesh will believe an upgrade landed before it took effect.
|
||||
|
||||
## References
|
||||
|
||||
|
||||
@@ -0,0 +1,310 @@
|
||||
---
|
||||
layer: to-be
|
||||
status: designed
|
||||
code: []
|
||||
updated: 2026-08-27
|
||||
decisions:
|
||||
- 02-DECISIONS/0036-a-node-is-a-managed-machine.md
|
||||
- 02-DECISIONS/0037-the-host-applies-it-does-not-decide.md
|
||||
- 02-DECISIONS/0038-a-node-joins-by-linking-first.md
|
||||
- 02-DECISIONS/0039-the-link-is-the-security-boundary.md
|
||||
- 02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md
|
||||
- 02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md
|
||||
---
|
||||
|
||||
# The node lifecycle
|
||||
|
||||
How a Linux machine becomes a node, stays one, and stops being one.
|
||||
|
||||
[`05-the-node-host.md`](05-the-node-host.md) describes the host as a component. This describes
|
||||
it as something that runs for years on a machine somebody else also uses — which is where the
|
||||
questions that were not being asked live.
|
||||
|
||||
## The states
|
||||
|
||||
```
|
||||
unmanaged ──install──► hosted ──enrol──► enrolled ⇄ disconnected
|
||||
▲ │
|
||||
└─────release──────┘
|
||||
```
|
||||
|
||||
| State | Has | Can |
|
||||
|---|---|---|
|
||||
| **unmanaged** | nothing of ours | — it is a Linux machine |
|
||||
| **hosted** | the host, no identity | apply a local file, apply its bundle |
|
||||
| **enrolled** | identity, link, store | everything; this is *a node* |
|
||||
| **disconnected** | identity, store, no link | hold its machine in the last state it was told |
|
||||
|
||||
**Only `enrolled` and `disconnected` are nodes**, and they are the same node in two situations
|
||||
rather than two kinds of thing
|
||||
([ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)). **`hosted` is not a
|
||||
node** — it is a machine with a program on it that has not been told which mesh it belongs to.
|
||||
|
||||
There is no state for *the first node*. That is the point of
|
||||
[ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md): the first node walks the
|
||||
same path, in an unusual order.
|
||||
|
||||
---
|
||||
|
||||
## unmanaged → hosted: installing
|
||||
|
||||
```
|
||||
pacman -S nox-mesh-host
|
||||
systemctl enable --now nox-mesh-host
|
||||
```
|
||||
|
||||
Or, where there is no repository to install from:
|
||||
|
||||
```
|
||||
curl -fsSL https://<release>/mesh-host-<version>-x86_64.tar.gz | tar -xz -C /usr/local/bin
|
||||
```
|
||||
|
||||
**The tarball must never acquire a dependency**, because the mesh's own package repository is
|
||||
hosted on the mesh. Any route that needs the mesh in order to install the thing that joins the
|
||||
mesh is a circle — unusable on a first node, and unusable by whoever is repairing a mesh that is
|
||||
down, which is exactly when it is wanted.
|
||||
|
||||
At this point the host is running and **doing nothing**. It has no identity, so there is nobody
|
||||
to link to and nothing to apply. It answers `profile`, `inventory` and `version`, and waits.
|
||||
|
||||
---
|
||||
|
||||
## hosted → enrolled: the ordinary case
|
||||
|
||||
```
|
||||
nox-mesh-host enrol --token <one-time token>
|
||||
```
|
||||
|
||||
The token carries three things and is carried by a person
|
||||
([ADR 0051](../../02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md)): the broker's
|
||||
address, the fingerprint to expect, and the right to join once.
|
||||
|
||||
What happens, in order:
|
||||
|
||||
1. the host dials the broker at the address in the token, **over the underlay**;
|
||||
2. it checks the broker's certificate against the pinned fingerprint — *before* sending anything;
|
||||
3. it presents the one-time secret and receives its **own durable identity**;
|
||||
4. it reports its `profile` and `inventory` upward;
|
||||
5. the control plane decides what this machine should be, and sends a declaration;
|
||||
6. the host applies it, reads back, and reports.
|
||||
|
||||
**Step 4 is the one that is easy to miss and is what makes step 5 possible.** The control plane
|
||||
cannot decide what a machine should run without knowing what it *can* run — a graphical session,
|
||||
a container runtime, an architecture. The profile is not a diagnostic; it is the input.
|
||||
|
||||
**The node computes nothing about the mesh.** It needs one peer to reach; the whole overlay is
|
||||
derived centrally and pushed down
|
||||
([ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md),
|
||||
[`08-connectivity.md`](08-connectivity.md)).
|
||||
|
||||
---
|
||||
|
||||
## hosted → enrolled: the first node
|
||||
|
||||
The same path, with the mesh built in the middle of it.
|
||||
|
||||
```
|
||||
# 1 — raise the substrate and the control plane from the carried bundle
|
||||
nox-mesh-host reconcile
|
||||
|
||||
# 2 — the control plane now exists, and issues the first token
|
||||
mesh-control token issue
|
||||
|
||||
# 3 — the machine joins the mesh it just raised
|
||||
nox-mesh-host enrol --token <token>
|
||||
```
|
||||
|
||||
Step 1 is the bootstrap from [`07-the-substrate.md`](07-the-substrate.md): Docker, then
|
||||
PostgreSQL, then the database, then the schema, then the control plane. It needs no identity
|
||||
because nothing is being asked of anyone — the host is applying a declaration it already
|
||||
carries, to the machine it is already on.
|
||||
|
||||
**After step 3 the first node is not special in any way**, which is the property `adopt.sh` and
|
||||
the bootstrap script never had. Its specialness lasted two commands.
|
||||
|
||||
**And enrolment is exercised on node one.** The path every other node depends on is walked
|
||||
immediately, against a control plane on the same machine, rather than being written and first
|
||||
used months later on node two.
|
||||
|
||||
---
|
||||
|
||||
## Adoption: what happens to what is already there
|
||||
|
||||
Adoption is not a state. It is what the **first apply** does when it is told to own something a
|
||||
machine already has ([research 012](../../01-RESEARCH/012-the-minimum-viable-node/00-overview.md)).
|
||||
|
||||
A candidate machine is not empty. It has a package manager, probably a container runtime,
|
||||
configuration somebody chose. [ADR 0043](../../02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md)
|
||||
says the host never touches what it did not create — adoption is the deliberate act of taking
|
||||
ownership of exactly that, so it is a companion to that rule rather than an exception:
|
||||
|
||||
> *never, unless adoption made it the host's* — with adoption **explicit, recorded, and visible
|
||||
> in what the host says it owns.**
|
||||
|
||||
Three rules, all earned:
|
||||
|
||||
**The original is kept before anything is written.** A one-way door on a working machine is not
|
||||
an installation. This is a *never* rule, and it earns that from the worst loss in this record —
|
||||
a tool acting on a path it did not own.
|
||||
|
||||
**On conflict, the machine's configuration wins.** Adoption always completes; the conflict is
|
||||
flagged and reconciled afterwards. A machine in use keeps working exactly as it did.
|
||||
|
||||
**Adoption produces a briefing**, not just a result: what it found, what it took over, and what
|
||||
it could not resolve — with each line marked `ok`, `kept`, `unknown` or `failed`, and the overall
|
||||
outcome **derived** from the worst line rather than stated alongside it.
|
||||
|
||||
---
|
||||
|
||||
## enrolled: what running actually looks like
|
||||
|
||||
Four reconcile triggers
|
||||
([ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md)):
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **start** | the machine may have changed while nothing was running |
|
||||
| **a declaration arrives** | the ordinary path |
|
||||
| **every ten minutes** | drift — something other than the host changed the machine |
|
||||
| **reconnect** | declarations may have been missed |
|
||||
|
||||
**Rebooting mid-apply is safe, and it is safe by construction.** The store records each resource
|
||||
*after* it worked ([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)), so
|
||||
a host that dies half way through comes back, finds the completed ones already matching, and
|
||||
applies the rest. The rule that exists to stop the host lying about what it did also makes it
|
||||
crash-safe.
|
||||
|
||||
---
|
||||
|
||||
## Updating what the node holds
|
||||
|
||||
An ordinary declaration. Someone assigns a module; the control plane recomputes what that node
|
||||
should be and sends it; the host applies the difference and removes what is no longer declared.
|
||||
|
||||
**Removal is not symmetric, and the asymmetry is the design:**
|
||||
|
||||
| | on being undeclared |
|
||||
|---|---|
|
||||
| file, directory | **removed** |
|
||||
| container | **removed** — the host created it |
|
||||
| service | **stopped**; the unit file is not the host's to delete |
|
||||
| package | **left installed** — *forgotten*, not removed |
|
||||
| action | **forgotten** — it left nothing the host owns |
|
||||
|
||||
The host removes what it *made* and leaves what it merely *configured*. Uninstalling a container
|
||||
runtime because a declaration changed would stop every container on the node.
|
||||
|
||||
---
|
||||
|
||||
## Updating the host itself
|
||||
|
||||
A declaration too — `package: nox-mesh-host`, at a version. The package manager writes the new
|
||||
binary; the running process is undisturbed, because Unix keeps the running executable's inode.
|
||||
|
||||
**Then the host exits, and the supervisor restarts it on the new binary.** After the apply
|
||||
completes, never during it; only when the executable actually changed; exit zero, so a restart
|
||||
is what happens next rather than a failure a supervisor backs off from.
|
||||
|
||||
**The host never asks the service manager to restart it.** That is the host killing itself
|
||||
part-way through an apply. It stops by finishing.
|
||||
|
||||
**A version-skewed fleet is normal**, because each node restarts when its own apply finishes.
|
||||
What a node reports must be the **running** version, not the installed one, or the mesh will
|
||||
believe an upgrade landed before it took effect.
|
||||
|
||||
---
|
||||
|
||||
## enrolled ⇄ disconnected
|
||||
|
||||
Not a failure. Not degraded. A situation
|
||||
([ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)).
|
||||
|
||||
A disconnected node **keeps reconciling against its own store**, so it goes on holding its
|
||||
machine in the last state it was told to hold. A laptop shut for a week comes back and
|
||||
reconciles; it does not come back and ask what it is.
|
||||
|
||||
What it cannot do: receive new declarations, be granted anything new, or have its certificates
|
||||
renewed — which is the clock on the whole arrangement
|
||||
([ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md)).
|
||||
|
||||
**How long it has been disconnected is a fact the mesh must hold**, and nothing holds it today.
|
||||
Without it, a node running last month's assignments looks exactly like one that is current.
|
||||
|
||||
---
|
||||
|
||||
## Rescue
|
||||
|
||||
The host is still a command-line tool, and that is what rescue is:
|
||||
|
||||
```
|
||||
nox-mesh-host owned # what do you think you own?
|
||||
nox-mesh-host apply repair.json # apply something by hand, locally
|
||||
nox-mesh-host profile # what can this machine actually do?
|
||||
```
|
||||
|
||||
`apply FILE` accepts actions, because someone who can write that file and run this binary as
|
||||
root can already do anything it can. The bound in
|
||||
[ADR 0047](../../02-DECISIONS/0047-the-bundle-may-carry-actions-the-link-may-not.md) is on what a
|
||||
**remote** party may push, not on what a person at the machine may do.
|
||||
|
||||
This replaces the three hand-run scripts that exist today — first node, joining, rescue — with
|
||||
one binary that has always been the same binary.
|
||||
|
||||
---
|
||||
|
||||
## enrolled → hosted: retiring a node
|
||||
|
||||
Two cases, and they are genuinely different.
|
||||
|
||||
**Graceful.** The control plane sends a final declaration that names nothing. The host removes
|
||||
what it owns by the table above, reports, and drops its identity. The machine keeps the host
|
||||
installed and is back to `hosted`. Nothing is left behind that anybody has to remember.
|
||||
|
||||
**The node is gone.** Stolen, dead, or simply unreachable. The mesh cannot tell it anything, and
|
||||
by [ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md) it will go on reconciling
|
||||
its last declaration **forever**.
|
||||
|
||||
That is the honest consequence of making disconnection ordinary, and the answer is not to make
|
||||
the host expire. It is that **the node holds nothing that outlives revocation**: its identity is
|
||||
its own, and every grant it holds is a per-node credential at the provider
|
||||
([ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md),
|
||||
[ADR 0045](../../02-DECISIONS/0045-a-context-owns-its-store.md)). Revoking is done at the
|
||||
database, the broker, the object store — not on the machine.
|
||||
|
||||
So a lost node keeps *running* and stops being able to *reach* anything. That is the best
|
||||
available outcome and it is worth stating plainly rather than implying the mesh can reach out and
|
||||
switch a machine off, which it cannot and should not be able to.
|
||||
|
||||
---
|
||||
|
||||
## Losing the store
|
||||
|
||||
Worth its own section because the failure is quiet.
|
||||
|
||||
If `/var/lib/mesh-host/state.json` is lost — a reinstall, a replaced disk — the host loses
|
||||
**its record of what it owns**, not its ability to work. It re-enrols, receives the declaration
|
||||
again, and re-applies it.
|
||||
|
||||
**What does not come back is removal.** Resources it applied under an older declaration and no
|
||||
longer holds a record of become unowned: the host will not touch them, because it never touches
|
||||
what it did not create. They sit there, unmanaged, indefinitely.
|
||||
|
||||
The store is therefore the one piece of node state that matters, and *how it is protected* is
|
||||
not designed.
|
||||
|
||||
---
|
||||
|
||||
## Open
|
||||
|
||||
- **Re-enrolling as the same node.** A machine that lost its identity gets a new token — but
|
||||
whether the mesh treats it as the same node or a new one is an operator's decision today, and
|
||||
nothing supports either.
|
||||
- **Protecting the store.** Above. Its loss is silent and permanent.
|
||||
- **How long disconnected, and who is told.** The fact is not held anywhere.
|
||||
- **Whether a `failed` adoption line still lets adoption complete.** Reopened by
|
||||
[research 012](../../01-RESEARCH/012-the-minimum-viable-node/00-overview.md) and not decided:
|
||||
*flags inform* was decided about conflicts, and a failure is different in kind.
|
||||
- **What a briefing looks like.** It is the first thing a session on a new node sees, which
|
||||
makes it an interface rather than a log.
|
||||
- **Where the enrolment token comes from, operationally.** A person carries it. Nothing says how
|
||||
it is generated, shown, or transported, and it is now the only secret in adoption.
|
||||
@@ -18,6 +18,7 @@ document is written and this one's status becomes `implemented`.
|
||||
| [`06-the-control-plane.md`](06-the-control-plane.md) | Tier 2 — what the term means, and the test for what belongs in it | [ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md) |
|
||||
| [`07-the-substrate.md`](07-the-substrate.md) | Tier 1 — what the control plane consumes and cannot grant itself | [ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md), [0048](../../02-DECISIONS/0048-the-substrate-is-named.md) |
|
||||
| [`08-connectivity.md`](08-connectivity.md) | One context in full — overlay, resolution, exposure, filtering, certificates | [ADR 0049](../../02-DECISIONS/0049-a-route-is-a-grant.md), [0050](../../02-DECISIONS/0050-reachability-is-a-property-of-the-address.md), [0051](../../02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md), [0055](../../02-DECISIONS/0055-the-control-plane-is-the-node-coordinating-contexts.md) |
|
||||
| [`09-the-node-lifecycle.md`](09-the-node-lifecycle.md) | How a machine becomes a node, stays one, and stops being one | [ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md), [0051](../../02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md) |
|
||||
|
||||
## Not yet written
|
||||
|
||||
|
||||
Reference in New Issue
Block a user