Settles the design repository now that the self-upgrade build is on main: - Records the two decisions that shipped without a record — ADR 0077 (the controller/foundation/node vocabulary) and ADR 0078 (the store and broker are ordinary modules); accepts ADR 0075 and 0076, which shipped work rests on. - Fills issue 051's amended-design and wires ADR 0078 into 07-the-foundation. - Sweeps the repo rename (mesh-control -> mesh-controller) into the mutable docs now that the forge repo is renamed; updates the glossary note and repos.md. - Fixes the six broken links from the design-doc renames, indexes the glossary, regenerates the decisions reading order. Both checks (records.py, index.py) are green. Statuses stay honest: the build is on main and lab-proven but not deployed as the production mesh, so the to-be docs remain in-progress and the as-is layer (the hal mesh) is unchanged — graduation to implemented + as-is belongs to deployment, not merge. https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
697 lines
33 KiB
Markdown
697 lines
33 KiB
Markdown
---
|
|
layer: to-be
|
|
status: in-progress
|
|
code:
|
|
- mesh-host internal/link/run.go
|
|
- mesh-host internal/link/enrol.go
|
|
- mesh-host packaging/nox-mesh-host-resume.service
|
|
- mesh-host packaging/nox-mesh-host-network.sh
|
|
- mesh-controller internal/token
|
|
- mesh-controller internal/inventory/nodes.go
|
|
updated: 2026-08-31
|
|
decisions:
|
|
- 02-DECISIONS/0004-a-node-and-how-it-joins.md
|
|
- 02-DECISIONS/0005-the-node-host.md
|
|
- 02-DECISIONS/0004-a-node-and-how-it-joins.md
|
|
- 02-DECISIONS/0004-a-node-and-how-it-joins.md
|
|
- 02-DECISIONS/0005-the-node-host.md
|
|
- 02-DECISIONS/0004-a-node-and-how-it-joins.md
|
|
- 02-DECISIONS/0005-the-node-host.md
|
|
- 02-DECISIONS/0010-delivery.md
|
|
- 02-DECISIONS/0005-the-node-host.md
|
|
- 02-DECISIONS/0005-the-node-host.md
|
|
- 02-DECISIONS/0005-the-node-host.md
|
|
---
|
|
|
|
# The node lifecycle
|
|
|
|
How a Linux machine becomes a node, stays one, and stops being one.
|
|
|
|
[`05-the-node-host.md`](05-the-node-host.md) describes the host as a component. This describes
|
|
it as something that runs for years on a machine somebody else also uses — which is where the
|
|
questions that were not being asked live.
|
|
|
|
## The states
|
|
|
|
```
|
|
unmanaged ──install──► hosted ──enrol──► enrolled ⇄ disconnected
|
|
▲ │
|
|
└─────release──────┘
|
|
```
|
|
|
|
| State | Has | Can |
|
|
|---|---|---|
|
|
| **unmanaged** | nothing of ours | — it is a Linux machine |
|
|
| **hosted** | the host, no identity | apply a local file, apply its bundle |
|
|
| **enrolled** | identity, link, store | everything; this is *a node* |
|
|
| **disconnected** | identity, store, no link | hold its machine in the last state it was told |
|
|
|
|
**Only `enrolled` and `disconnected` are nodes**, and they are the same node in two situations
|
|
rather than two kinds of thing
|
|
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). **`hosted` is not a
|
|
node** — it is a machine with a program on it that has not been told which mesh it belongs to.
|
|
|
|
There is no state for *the first node*. That is the point of
|
|
[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md): the first node walks the
|
|
same path, in an unusual order.
|
|
|
|
---
|
|
|
|
## unmanaged → hosted: installing
|
|
|
|
In the machine's own idiom, because the package manager and the init file are the system's
|
|
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)):
|
|
|
|
```
|
|
# Alpine — the intended first node
|
|
apk add nox-mesh-host
|
|
rc-update add nox-mesh-host && rc-service nox-mesh-host start
|
|
|
|
# Arch
|
|
pacman -S nox-mesh-host
|
|
systemctl enable --now nox-mesh-host
|
|
```
|
|
|
|
Two lines each, and the init file behind them is four
|
|
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)) — it says
|
|
*run the launcher at boot* and nothing else, so a third system is transcription rather than a
|
|
port.
|
|
|
|
Or, where there is no repository to install from:
|
|
|
|
```
|
|
curl -fsSL https://<release>/mesh-host-<system>-<version>-x86_64.tar.gz | tar -xz -C /usr/local/bin
|
|
```
|
|
|
|
**The binary is per system as well as per architecture**, because two of its appliers are.
|
|
|
|
**The tarball must never acquire a dependency**, because the mesh's own package repository is
|
|
hosted on the mesh. Any route that needs the mesh in order to install the thing that joins the
|
|
mesh is a circle — unusable on a first node, and unusable by whoever is repairing a mesh that is
|
|
down, which is exactly when it is wanted.
|
|
|
|
### The unit it installs
|
|
|
|
```ini
|
|
[Unit]
|
|
Description=Novox Mesh node host
|
|
After=network-online.target
|
|
Wants=network-online.target
|
|
|
|
[Service]
|
|
ExecStart=/usr/lib/nox-mesh-host/launch
|
|
Restart=always
|
|
RestartSec=5s
|
|
StateDirectory=mesh-host
|
|
|
|
[Install]
|
|
WantedBy=multi-user.target
|
|
```
|
|
|
|
**Two lines of policy, and that is deliberate**
|
|
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)). The init is
|
|
asked to *start this at boot* and *start it again if it exits*, and nothing else. Both are
|
|
expressible in OpenRC, runit, s6 and an Android `init.rc`, so porting this file is transcription
|
|
rather than design.
|
|
|
|
**`Restart=always` and not `on-failure`**: the host restarts onto a new binary by exiting
|
|
*cleanly*, so a supervisor that only restarts on failure would leave every upgraded node stopped,
|
|
having successfully upgraded.
|
|
|
|
**What the init does not do is decide when to give up.** Counting failed starts and rolling back
|
|
lives in the launcher, where it can be tested — `OnFailure=` in a unit file can only be read and
|
|
hoped for, and it is the one thing that has to work on a machine where nothing else does.
|
|
|
|
**The package owns this file. The host never does.** It manages `service` resources, and its own
|
|
unit is a service — the temptation is obvious and it ends with a host stopping itself half way
|
|
through an apply, leaving a machine with nothing running to fix it. A declaration naming the
|
|
host's own unit is **refused**, and that refusal is a test rather than a convention.
|
|
|
|
The line to hold: **the installation owns the host; the host owns everything else.**
|
|
|
|
At this point the host is running and **doing nothing**. It has no identity, so there is nobody
|
|
to link to and nothing to apply. It answers `profile`, `inventory` and `version`, and waits.
|
|
|
|
---
|
|
|
|
## hosted → enrolled: the ordinary case
|
|
|
|
```
|
|
nox-mesh-host enrol --token <one-time token>
|
|
```
|
|
|
|
The token carries **four** things and is carried by a person
|
|
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)): the broker's
|
|
address, the fingerprint to expect, **the controller's signing identity**, and the right to
|
|
join once.
|
|
|
|
**The fourth is the one this document listed three of.** A node connects to the broker and takes
|
|
instruction from the controller behind it, and those are two different identities. Pinning only
|
|
the broker would make the controller's authority *transitive* — a compromised broker could then
|
|
forge declarations, which, since the host applies whatever the link delivers, is the whole machine.
|
|
So the transport is verified once at connect, and **each declaration is verified by its signature,
|
|
every time**.
|
|
|
|
What happens, in order:
|
|
|
|
1. the host dials the broker at the address in the token, **over the underlay**;
|
|
2. it checks the broker's certificate against the pinned fingerprint — *before* sending anything;
|
|
3. it presents the one-time secret **and its own public key**, which the mesh records;
|
|
4. it reports its `profile` and `inventory` upward;
|
|
5. the controller decides what this machine should be, and sends a declaration;
|
|
6. the host applies it, reads back, and reports.
|
|
|
|
**Step 4 is the one that is easy to miss and is what makes step 5 possible.** The controller
|
|
cannot decide what a machine should run without knowing what it *can* run — a graphical session,
|
|
a container runtime, an architecture. The profile is not a diagnostic; it is the input.
|
|
|
|
**The node computes nothing about the mesh.** It needs one peer to reach; the whole overlay is
|
|
derived centrally and pushed down
|
|
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md),
|
|
[`08-connectivity.md`](08-connectivity.md)).
|
|
|
|
### The first declaration is the overlay, and nothing else
|
|
|
|
**The mesh makes a node reachable before it makes it useful.** Step 5 is not one declaration
|
|
carrying everything the node will ever run. It is two, in order:
|
|
|
|
```
|
|
first the overlay — this node's address, its keys, its peers, its names
|
|
then everything else — packages, containers, services, files
|
|
```
|
|
|
|
Three reasons, and the third is the one that matters when something goes wrong:
|
|
|
|
- **It is forced.** A node cannot join the overlay before contacting the mesh, because its
|
|
address and peer set are *assigned* — it generates a keypair, publishes the public half, and
|
|
receives the rest ([`08-connectivity.md`](08-connectivity.md)). So the overlay is the first
|
|
thing the mesh can give it, and it should be.
|
|
- **It is what [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) already
|
|
says:** *a joining node does the minimum to be reachable, and nothing else.*
|
|
- **It is the way back in.** Once the overlay is up, the node is reachable over it — by SSH, by
|
|
anything. If a later declaration breaks the machine, there is a route to it that does not
|
|
depend on the mesh's control path working. **Sending a large first declaration risks a node
|
|
that is broken and unreachable at the same time**, and those two failures are much worse
|
|
together than separately.
|
|
|
|
### Reachable is not the same as having a control surface
|
|
|
|
Worth stating plainly, because the two rules read as a contradiction and are not.
|
|
|
|
| | |
|
|
|---|---|
|
|
| **every node reaches every other node** | over the overlay — SSH, services, ordinary traffic. This is the point of having one |
|
|
| **every node consumes from the broker** | its own queue, over its own outbound connection ([ADR 0002](../../02-DECISIONS/0002-nodes-communicate-over-a-broker.md)) |
|
|
| **nothing dials a node to control it** | the host has no inbound control surface ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) |
|
|
|
|
**[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) is about the control
|
|
channel, not about network reachability.** What it forbids is a listening thing that accepts
|
|
instructions and changes the machine. A node being reachable on the overlay — the whole purpose
|
|
of the overlay — is untouched by it, and so is a person opening a shell on it.
|
|
|
|
The distinction is *who can tell this machine what to be*: only the controller, only over the
|
|
link the node opened, only in declarations of known shape.
|
|
|
|
---
|
|
|
|
## hosted → enrolled: the first node
|
|
|
|
The same path, with the mesh built in the middle of it.
|
|
|
|
```
|
|
# 1 — raise the foundation and the controller from the carried bundle
|
|
nox-mesh-host reconcile
|
|
|
|
# 2 — the controller now exists, and issues the first token
|
|
mesh-controller token issue
|
|
|
|
# 3 — the machine joins the mesh it just raised
|
|
nox-mesh-host enrol --token <token>
|
|
```
|
|
|
|
Step 1 is the bootstrap from [`07-the-foundation.md`](07-the-foundation.md): a container runtime,
|
|
then PostgreSQL, then the database, then the schema, then the controller. It needs no identity
|
|
because nothing is being asked of anyone — the host is applying a declaration it already
|
|
carries, to the machine it is already on.
|
|
|
|
**After step 3 the first node is not special in any way**, which is the property `adopt.sh` and
|
|
the bootstrap script never had. Its specialness lasted two commands.
|
|
|
|
**And enrolment is exercised on node one.** The path every other node depends on is walked
|
|
immediately, against a controller on the same machine, rather than being written and first
|
|
used months later on node two.
|
|
|
|
---
|
|
|
|
## Two kinds of host
|
|
|
|
Everything above assumes a machine with an init that runs the host at boot. Not every machine
|
|
has one ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)).
|
|
|
|
| | **resident** | **episodic** |
|
|
|---|---|---|
|
|
| examples | Alpine, Arch | Android |
|
|
| started by | an init, at boot | whatever the platform allows |
|
|
| supervised by | the launcher | nothing — the platform decides when it runs |
|
|
| the link | held open | opened while it runs |
|
|
| being stopped | shutdown, or a failure | **ordinary** |
|
|
| shapes | all six | `file`, `directory`, `action` |
|
|
| can be the first node | yes | **no** |
|
|
|
|
**An episodic host being killed is disconnection, not failure.** That is
|
|
[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) doing the work it was written
|
|
for: reachability is state, not class. Everything the design already does for a laptop that
|
|
closes — an authoritative local store, reconcile on start, *last heard from* reported without an
|
|
alarm — is what an episodic host needs, at a shorter period.
|
|
|
|
**It cannot be the first node**, and that is not a limitation to work around. Every step of
|
|
raising a foundation is a `package`, a `container` or an `action` against one, and a partial host
|
|
refuses the first two. So `mesh-host-android bundle` returns a file that says so rather than an
|
|
empty placeholder waiting to be filled in.
|
|
|
|
**Two things this changes for anything reading the mesh.** *Last heard from* is a much weaker
|
|
signal on an episodic host — a healthy phone looks like a dead server — so a reader has to know
|
|
which kind it is looking at. And a declaration may take a long time to land, which makes
|
|
[ADR 0010](../../02-DECISIONS/0010-delivery.md)'s separation of
|
|
*outstanding* from *failed* load-bearing rather than tidy.
|
|
|
|
**Still open:** how an episodic host is started in practice — an APK with a foreground service,
|
|
or Termux with its boot addon — and, first, **what an Android node is for.** A device that can
|
|
write files and run commands is not a workload host; it is a presence, or somewhere an agent
|
|
runs. Building the start mechanism before deciding that would be building it for nobody.
|
|
|
|
## Adoption: what happens to what is already there
|
|
|
|
Adoption is not a state. It is what the **first apply** does when it is told to own something a
|
|
machine already has ([research 012](../../01-RESEARCH/012-the-minimum-viable-node/00-overview.md)).
|
|
|
|
A candidate machine is not empty. It has a package manager, probably a container runtime,
|
|
configuration somebody chose. [ADR 0005](../../02-DECISIONS/0005-the-node-host.md)
|
|
says the host never touches what it did not create — adoption is the deliberate act of taking
|
|
ownership of exactly that, so it is a companion to that rule rather than an exception:
|
|
|
|
> *never, unless adoption made it the host's* — with adoption **explicit, recorded, and visible
|
|
> in what the host says it owns.**
|
|
|
|
Three rules, all earned:
|
|
|
|
**The original is kept before anything is written.** A one-way door on a working machine is not
|
|
an installation. This is a *never* rule, and it earns that from the worst loss in this record —
|
|
a tool acting on a path it did not own.
|
|
|
|
**On conflict, the machine's configuration wins.** Adoption always completes; the conflict is
|
|
flagged and reconciled afterwards. A machine in use keeps working exactly as it did.
|
|
|
|
**Adoption produces a briefing**, not just a result: what it found, what it took over, and what
|
|
it could not resolve — with each line marked `ok`, `kept`, `unknown` or `failed`, and the overall
|
|
outcome **derived** from the worst line rather than stated alongside it.
|
|
|
|
---
|
|
|
|
## enrolled: what running actually looks like
|
|
|
|
**Changes are pushed, not polled.** A declaration arrives as a message on the link and the host
|
|
applies it then. The link is already open and outbound
|
|
([ADR 0002](../../02-DECISIONS/0002-nodes-communicate-over-a-broker.md),
|
|
[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) — asking it
|
|
repeatedly whether anything has changed would be slower to land *and* constant traffic to learn
|
|
nothing.
|
|
|
|
| Trigger | Kind | |
|
|
|---|---|---|
|
|
| **a declaration arrives** | **pushed** | the ordinary path — this is how a change lands |
|
|
| **start** | event | the machine may have changed while nothing was running |
|
|
| **reconnect** | event | declarations may have been missed |
|
|
| **every ten minutes** | periodic | **drift, and only drift** |
|
|
|
|
**Why the timer cannot be an event.** Drift is change the *mesh did not make* — somebody edited
|
|
a managed file, a distribution upgrade replaced a config, a container was stopped by hand.
|
|
Nothing will ever publish a message about it, because whatever did it is not part of the mesh.
|
|
Only looking finds it.
|
|
|
|
So the two periodic things do different jobs and should not be conflated:
|
|
|
|
| | direction | answers |
|
|
|---|---|---|
|
|
| **reconcile timer** | local, looks at the machine | *does this machine still match what it was told?* |
|
|
| **heartbeat** | upward, reports to the mesh | *is this node still here, and what is it running?* |
|
|
|
|
**The heartbeat is what makes silence mean something.** A node with nothing to do sends nothing;
|
|
without a heartbeat that is indistinguishable from a node that stopped. With one, *last heard
|
|
from* is a fact beside every node — which is what
|
|
[how long disconnected](#how-long-disconnected-and-who-is-told) reports and what
|
|
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) exists
|
|
because a stuck node cannot send.
|
|
|
|
**Rebooting mid-apply is safe by construction.** The store records each resource *after* it
|
|
worked ([ADR 0018](../../02-DECISIONS/0018-a-picture-is-read-from-what-runs.md)), so a host that
|
|
dies half way through comes back, finds the completed ones already matching, and applies the
|
|
rest. The rule that exists to stop the host lying about what it did also makes it crash-safe.
|
|
|
|
## Updating what the node holds
|
|
|
|
An ordinary declaration. Someone assigns a module; the controller recomputes what that node
|
|
should be and sends it; the host applies the difference and removes what is no longer declared.
|
|
|
|
**Removal is not symmetric, and the asymmetry is the design:**
|
|
|
|
| | on being undeclared |
|
|
|---|---|
|
|
| file, directory | **removed** |
|
|
| container | **removed** — the host created it |
|
|
| service | **stopped**; the unit file is not the host's to delete |
|
|
| package | **left installed** — *forgotten*, not removed |
|
|
| action | **forgotten** — it left nothing the host owns |
|
|
|
|
The host removes what it *made* and leaves what it merely *configured*. Uninstalling a container
|
|
runtime because a declaration changed would stop every container on the node.
|
|
|
|
---
|
|
|
|
## enrolled ⇄ disconnected
|
|
|
|
Not a failure. Not degraded. A situation
|
|
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)).
|
|
|
|
A disconnected node **keeps reconciling against its own store**, so it goes on holding its
|
|
machine in the last state it was told to hold. A laptop shut for a week comes back and
|
|
reconciles; it does not come back and ask what it is.
|
|
|
|
What it cannot do: receive new declarations, be granted anything new, or have its certificates
|
|
renewed — which is the clock on the whole arrangement
|
|
([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)).
|
|
|
|
**How long it has been disconnected is a fact the mesh must hold**, and nothing holds it today.
|
|
Without it, a node running last month's assignments looks exactly like one that is current.
|
|
|
|
---
|
|
|
|
## Rescue
|
|
|
|
The host is still a command-line tool, and that is what rescue is:
|
|
|
|
```
|
|
nox-mesh-host owned # what do you think you own?
|
|
nox-mesh-host apply repair.json # apply something by hand, locally
|
|
nox-mesh-host profile # what can this machine actually do?
|
|
```
|
|
|
|
`apply FILE` accepts actions, because someone who can write that file and run this binary as
|
|
root can already do anything it can. The bound in
|
|
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) is on what a
|
|
**remote** party may push, not on what a person at the machine may do.
|
|
|
|
This replaces the three hand-run scripts that exist today — first node, joining, rescue — with
|
|
one binary that has always been the same binary.
|
|
|
|
---
|
|
|
|
## enrolled → hosted: retiring a node
|
|
|
|
Two cases, and they are genuinely different.
|
|
|
|
**Graceful.** The controller sends a final declaration that names nothing. The host removes
|
|
what it owns by the table above, reports, and drops its identity. The machine keeps the host
|
|
installed and is back to `hosted`. Nothing is left behind that anybody has to remember.
|
|
|
|
**The node is gone.** Stolen, dead, or simply unreachable. The mesh cannot tell it anything, and
|
|
by [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) it will go on reconciling
|
|
its last declaration **forever**.
|
|
|
|
That is the honest consequence of making disconnection ordinary, and the answer is not to make
|
|
the host expire. It is that **the node holds nothing that outlives revocation**: its identity is
|
|
its own, and every grant it holds is a per-node credential at the provider
|
|
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md),
|
|
[ADR 0008](../../02-DECISIONS/0008-a-context-owns-its-store.md)). Revoking is done at the
|
|
database, the broker, the object store — not on the machine.
|
|
|
|
So a lost node keeps *running* and stops being able to *reach* anything. That is the best
|
|
available outcome and it is worth stating plainly rather than implying the mesh can reach out and
|
|
switch a machine off, which it cannot and should not be able to.
|
|
|
|
---
|
|
|
|
## Losing the store
|
|
|
|
Worth its own section because the failure is quiet.
|
|
|
|
If `/var/lib/mesh-host/state.json` is lost — a reinstall, a replaced disk — the host loses
|
|
**its record of what it owns**, not its ability to work. It re-enrols, receives the declaration
|
|
again, and re-applies it.
|
|
|
|
**Without help, what does not come back is removal.** Resources applied under an older
|
|
declaration, whose record is gone, become unowned: the host will not touch them, because it
|
|
never touches what it did not create. They would sit there, unmanaged, indefinitely.
|
|
|
|
**So the mesh keeps a copy of what each node reports it owns**, refreshed on every apply report,
|
|
and hands it back on a rebuild — see [Protecting the store](#protecting-the-store). The store
|
|
remains locally authoritative for *operating*; the copy exists only for this.
|
|
|
|
---
|
|
|
|
## Upgrading the host
|
|
|
|
The host is delivered like anything else
|
|
([ADR 0010](../../02-DECISIONS/0010-delivery.md)), and this is worth
|
|
walking through because tier 0 looks like it should be special and is not.
|
|
|
|
```
|
|
push to mesh-host
|
|
│
|
|
├─ build go build → one static binary
|
|
├─ publish packaged, into the mesh's own package repository
|
|
└─ deploy each node's declaration now names the new version
|
|
│
|
|
└─ pushed to each node; the host applies it on arrival
|
|
(a node that is offline gets it on reconnect)
|
|
```
|
|
|
|
**Compared with today.** The current pipeline's third silo runs *once per node* and sends each
|
|
one a command to install and start. That is where the as-is records a package install that
|
|
404ed while the job went green. Here deploy is **one write** — the declaration changes — and the
|
|
installing is the host's ordinary work, which reads back before it records anything.
|
|
|
|
**The repository is reachable because a declaration made it so.** A `file` resource writes the
|
|
package manager's configuration pointing at the mesh's repository; a `package` resource names
|
|
the version. Both ordinary shapes, applied by the same host. **No new resource type**, which is
|
|
the test of whether this is really uniform.
|
|
|
|
### The restart
|
|
|
|
```
|
|
1 pacman installs the new binary the running process is untouched —
|
|
Unix keeps the running executable's inode
|
|
2 the host verifies the new binary runs `nox-mesh-host version`, as a subprocess
|
|
3 it finishes the apply and reports never mid-way
|
|
4 it exits 0 having finished, not having been stopped
|
|
5 the launcher starts it again on the new binary — it supervises the host
|
|
rather than exec'ing it (ADR 0005), so this
|
|
needs nothing from the init
|
|
6 the new host reconciles on start trigger 1, confirming the machine still matches
|
|
```
|
|
|
|
**Step 2 is the one to insist on.** A package can install a binary that does not execute here —
|
|
wrong architecture, a libc that is not present. Running it once before committing to a restart
|
|
turns "the node never came back" into "the apply failed and said why". It is the same read-back
|
|
rule the rest of the host already follows, applied to the one resource that is the host.
|
|
|
|
**The host never asks the service manager to restart it.** That is the host stopping itself
|
|
part-way through an apply. It stops by finishing.
|
|
|
|
**A fleet upgrades over an interval, not at an instant**, because each node restarts when its
|
|
own apply completes. A node must therefore report the version it is **running**, not the one
|
|
installed — otherwise the mesh believes an upgrade landed at step 1.
|
|
|
|
**A version that crashes on start rolls itself back**
|
|
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)).
|
|
|
|
What the init starts is not the host but a **launcher**, and the launcher is where the policy
|
|
lives:
|
|
|
|
```
|
|
init ──► nox-mesh-host-launch ──► nox-mesh-host
|
|
├─ halted? say so and stop; a person has to look
|
|
├─ count this start attempt
|
|
├─ too many, not yet rolled back? roll back, then start
|
|
├─ too many, already rolled back? halt — the machine is the problem
|
|
└─ otherwise start the host
|
|
```
|
|
|
|
It reinstalls the version recorded in `known-good`, which the host wrote the last time it
|
|
completed a reconcile — and the host clears the attempt counter at the same moment, for the same
|
|
reason.
|
|
|
|
**The launcher rather than the init's own features**, because this is the one thing that must
|
|
work on a machine where nothing else does. A shell script with a counter can be run against a
|
|
stub package manager and asserted; `OnFailure=` in a unit file can only be read and hoped for.
|
|
It also means the init is asked for nothing but *start* and *restart*, which every init can
|
|
do.
|
|
|
|
**It rolls back once.** If the previous version also fails, the node stops in a failed state
|
|
rather than flapping between two binaries. A second failure is a different diagnosis: the
|
|
machine is the problem, not the binary.
|
|
|
|
**Why this matters more than it looks.** A host that will not start cannot link, and a node that
|
|
is not linking looks exactly like a machine somebody switched off — which is the one condition
|
|
this design has deliberately decided not to alarm on. Without rollback, a bad release reaches
|
|
every node, each one goes quiet, and the mesh reports a fleet of sleeping laptops.
|
|
|
|
---
|
|
|
|
## Details that are easy to get wrong
|
|
|
|
Each of these has a wrong answer that looks reasonable, which is why they are written down
|
|
rather than left to be worked out.
|
|
|
|
### Re-enrolling as the same node
|
|
|
|
**A token is issued *for* a node record**, and that is where a re-enrolment is decided.
|
|
|
|
```
|
|
mesh-controller token issue --node workstation # this machine is that node again
|
|
mesh-controller token issue --new # a machine the mesh has not seen
|
|
```
|
|
|
|
The host does not need to know which it is. It presents a token and receives an identity; what
|
|
that identity is bound to was decided when the token was made.
|
|
|
|
**Issuing a re-enrolment token revokes the previous identity for that node**, and that is not
|
|
housekeeping. Two live identities for one node record is the stolen-laptop case with the thief's
|
|
credentials still valid — the case
|
|
[`retiring a node`](#enrolled--hosted-retiring-a-node) says is answered by revocation.
|
|
|
|
### Protecting the store
|
|
|
|
**The host reports what it owns, and the mesh keeps the last report.**
|
|
|
|
The store stays locally authoritative — a node operates from its own copy and needs nothing to
|
|
do so ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). What changes is that
|
|
the mesh holds a **copy for recovery**, refreshed on every apply report.
|
|
|
|
So a node that loses its state file re-enrols, receives both the declaration *and* the record of
|
|
what it previously owned, and can then remove what is no longer declared. The orphans that used
|
|
to be permanently stranded are recoverable.
|
|
|
|
**This is a backup, never a source.** The host never reads it to decide anything; it is handed
|
|
back only on a store rebuild, and a node that disagrees with it wins, because the node is the
|
|
one that can see the machine.
|
|
|
|
### How long disconnected, and who is told
|
|
|
|
**The mesh records last contact per node; the node records time since it last linked.** Both,
|
|
because they answer different questions — the mesh's is *have I heard from it*, the node's is
|
|
*how stale am I*, and a node reporting the second on reconnect is how a long absence gets
|
|
noticed at all.
|
|
|
|
**No threshold and no alarm.** A laptop switched off for three weeks is doing nothing wrong, and
|
|
a mesh that alerted on it would train people to ignore the alert. It is a **reported fact** —
|
|
`last seen 4 days ago` beside every node — and what counts as too long is a judgement for
|
|
whoever is looking, not a constant in the design.
|
|
|
|
### Whether a failed adoption line blocks
|
|
|
|
**Adoption always completes. A node with a `failed` line is a node, and it is not eligible for
|
|
assignment until the failure is resolved.**
|
|
|
|
*Flags inform, they do not block* holds for **conflicts** — where the mesh chose deliberately and
|
|
the machine still works. A **failure** is different in kind: not *we chose* but *we could not*,
|
|
and it gets different treatment for that reason.
|
|
|
|
The distinction is between **joining** and **being given work**. Refusing to join makes a
|
|
machine in use unadoptable, which is the outcome that rule exists to prevent. Placing work on a
|
|
machine where something the mesh needed never happened produces a module that is installed and
|
|
does not work — [04-ISSUES/007](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)
|
|
arriving from the adoption side.
|
|
|
|
### What a briefing is
|
|
|
|
**A structured document with prose in it**, held in the node's state and reported to the mesh.
|
|
It is the first thing a session on a new node reads, which makes it an interface.
|
|
|
|
```
|
|
outcome kept derived from the worst line below, never stated separately
|
|
node workstation
|
|
adopted 2026-08-27T14:02Z
|
|
|
|
ok container runtime docker 27.0, adopted; original config kept at <path>
|
|
kept storage driver machine has overlay2, the mesh wanted btrfs — machine wins
|
|
unknown firewall ruleset could not be parsed
|
|
failed package database locked by another process
|
|
|
|
what to look at
|
|
The storage driver disagreement is preference, not requirement, so nothing is broken.
|
|
The package database was locked; nothing was installed. Re-run adoption when it is free.
|
|
```
|
|
|
|
**The outcome is computed from the lines**, so a briefing cannot read *fine* while carrying a
|
|
failed line. Two independently written fields drift, and that drift is the fault this repository
|
|
keeps cataloguing.
|
|
|
|
### Where the enrolment token comes from
|
|
|
|
**`mesh-controller token issue` prints it once**, to the person running it. Single-use, and it
|
|
expires whether used or not ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)).
|
|
|
|
It is carried by hand — read off a screen, pasted into a terminal. That is the design rather than
|
|
a gap in it: its authenticity comes from the channel it travelled, which is what
|
|
lets a node verify a mesh it has never spoken to
|
|
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). A token emailed,
|
|
committed, or dropped in shared storage has lost the only property that makes it worth carrying.
|
|
|
|
**On the first node it comes from the controller that was raised two commands ago**, which is
|
|
the same command against a mesh that is one machine old.
|
|
|
|
---
|
|
|
|
## Still open
|
|
|
|
- ~~**Automatic rollback of a bad host version.**~~ **Resolved** by
|
|
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md): a launcher
|
|
counts failed starts and rolls back — shipped by the package, not the host binary, because a
|
|
binary that will not start cannot recover itself. It rolls back once; a second failure means
|
|
the machine is the problem, not the binary.
|
|
- **How a previous declaration is retained and chosen**, which is what rollback of anything else
|
|
would use ([ADR 0010](../../02-DECISIONS/0010-delivery.md)).
|
|
- **A node returning after months** applies a very large jump in one go. Correct, and untested.
|
|
|
|
## Going away and coming back
|
|
|
|
*2026-08-31, from being asked whether a machine that drops off has to be adopted again.*
|
|
|
|
**It does not, and nothing about it expires.** A disconnected node is the same node in a different
|
|
situation ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) — reachability is state,
|
|
not class. The node holds its own identity and the mesh holds the public half; there is no lease,
|
|
no timeout, and nothing that lapses while a machine is shut. A laptop closed for a week comes back
|
|
and reconnects, and a declaration sent while it was away is waiting for it.
|
|
|
|
**The only thing that forces re-enrolment is a machine losing its own key** — a reinstall, an image
|
|
re-cloned. That is deliberate: the mesh then reports what it can no longer seal to rather than
|
|
delivering blobs the machine cannot open.
|
|
|
|
### The gap that was left, and why it mattered
|
|
|
|
**A machine that suspends does not know it has been disconnected, and neither does anything else.**
|
|
After a resume the socket looks perfectly healthy from inside the process: no error, no close,
|
|
because nothing has tried to send anything yet. Heartbeats discover it twenty or thirty seconds
|
|
later.
|
|
|
|
For those twenty or thirty seconds the node believes it is in the mesh and is not — and *absence
|
|
must never be indistinguishable from a failure to answer* is the rule this whole design is built
|
|
on. It recovered on its own, which is why this was a quality gap rather than a fault. It was still
|
|
the machine waiting to be told something it already knew.
|
|
|
|
**So the machine says so.** Waking, and changing network, both rouse the host.
|
|
|
|
| | |
|
|
|---|---|
|
|
| **it ends the current attempt** | not the wait after it — the process is not waiting, it is sitting inside a connection that will never return |
|
|
| **by signal, not by anything listening** | a socket for this would be a control surface on every machine, in exchange for saving twenty seconds, and the security argument rests on there not being one |
|
|
| **two rouses at once are one** | a machine suspending and resuming repeatedly must not build a backlog of reconnections to work through |
|
|
| **the backoff is not reset** | being roused says the machine changed, not that whatever refused the connection has stopped. A laptop woken on a network with no route would otherwise retry at full speed for as long as somebody keeps opening the lid |
|
|
| **`down` does not rouse** | the link is already gone, reconnecting will fail, and the backoff exists for exactly that |
|
|
|
|
*Checked by holding a link that never returns on its own — which is precisely what a suspended
|
|
connection is — rousing it, and requiring the attempt to end and another to begin. And by running
|
|
the dispatcher against every event a network manager emits, requiring it to act on the ones that
|
|
change where packets go and on no others.*
|