Record that a machine coming back is ordinary, and what waking now does

Asked whether a machine that drops off needs re-adopting: it does not, nothing
expires, and the only thing that forces re-enrolment is losing its own key.

The gap was the twenty or thirty seconds after a resume in which a node
believes it is in a mesh it has left — recovering on its own, which made it a
quality gap rather than a fault, and still a machine waiting to be told
something it already knew.
This commit is contained in:
2026-08-31 10:21:50 +02:00
parent 6374c1eb60
commit 35ce23fb76
@@ -2,7 +2,10 @@
layer: to-be layer: to-be
status: designed status: designed
code: code:
- mesh-host internal/link/run.go
- mesh-host internal/link/enrol.go - mesh-host internal/link/enrol.go
- mesh-host packaging/nox-mesh-host-resume.service
- mesh-host packaging/nox-mesh-host-network.sh
- mesh-control internal/token - mesh-control internal/token
- mesh-control internal/inventory/nodes.go - mesh-control internal/inventory/nodes.go
updated: 2026-08-27 updated: 2026-08-27
@@ -650,3 +653,44 @@ the same command against a mesh that is one machine old.
- **How a previous declaration is retained and chosen**, which is what rollback of anything else - **How a previous declaration is retained and chosen**, which is what rollback of anything else
would use ([ADR 0010](../../02-DECISIONS/0010-delivery.md)). would use ([ADR 0010](../../02-DECISIONS/0010-delivery.md)).
- **A node returning after months** applies a very large jump in one go. Correct, and untested. - **A node returning after months** applies a very large jump in one go. Correct, and untested.
## Going away and coming back
*2026-08-31, from being asked whether a machine that drops off has to be adopted again.*
**It does not, and nothing about it expires.** A disconnected node is the same node in a different
situation ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) — reachability is state,
not class. The node holds its own identity and the mesh holds the public half; there is no lease,
no timeout, and nothing that lapses while a machine is shut. A laptop closed for a week comes back
and reconnects, and a declaration sent while it was away is waiting for it.
**The only thing that forces re-enrolment is a machine losing its own key** — a reinstall, an image
re-cloned. That is deliberate: the mesh then reports what it can no longer seal to rather than
delivering blobs the machine cannot open.
### The gap that was left, and why it mattered
**A machine that suspends does not know it has been disconnected, and neither does anything else.**
After a resume the socket looks perfectly healthy from inside the process: no error, no close,
because nothing has tried to send anything yet. Heartbeats discover it twenty or thirty seconds
later.
For those twenty or thirty seconds the node believes it is in the mesh and is not — and *absence
must never be indistinguishable from a failure to answer* is the rule this whole design is built
on. It recovered on its own, which is why this was a quality gap rather than a fault. It was still
the machine waiting to be told something it already knew.
**So the machine says so.** Waking, and changing network, both rouse the host.
| | |
|---|---|
| **it ends the current attempt** | not the wait after it — the process is not waiting, it is sitting inside a connection that will never return |
| **by signal, not by anything listening** | a socket for this would be a control surface on every machine, in exchange for saving twenty seconds, and the security argument rests on there not being one |
| **two rouses at once are one** | a machine suspending and resuming repeatedly must not build a backlog of reconnections to work through |
| **the backoff is not reset** | being roused says the machine changed, not that whatever refused the connection has stopped. A laptop woken on a network with no route would otherwise retry at full speed for as long as somebody keeps opening the lid |
| **`down` does not rouse** | the link is already gone, reconnecting will fail, and the backoff exists for exactly that |
*Checked by holding a link that never returns on its own — which is precisely what a suspended
connection is — rousing it, and requiring the attempt to end and another to begin. And by running
the dispatcher against every event a network manager emits, requiring it to act on the ones that
change where packets go and on no others.*