From 35ce23fb7654e78ce3f582a713a0b51a5b5a83c4 Mon Sep 17 00:00:00 2001 From: jochen Date: Mon, 31 Aug 2026 10:21:50 +0200 Subject: [PATCH] Record that a machine coming back is ordinary, and what waking now does MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Asked whether a machine that drops off needs re-adopting: it does not, nothing expires, and the only thing that forces re-enrolment is losing its own key. The gap was the twenty or thirty seconds after a resume in which a node believes it is in a mesh it has left — recovering on its own, which made it a quality gap rather than a fault, and still a machine waiting to be told something it already knew. --- 03-DESIGN/01-to-be/09-the-node-lifecycle.md | 44 +++++++++++++++++++++ 1 file changed, 44 insertions(+) diff --git a/03-DESIGN/01-to-be/09-the-node-lifecycle.md b/03-DESIGN/01-to-be/09-the-node-lifecycle.md index c797723..3bb36f2 100644 --- a/03-DESIGN/01-to-be/09-the-node-lifecycle.md +++ b/03-DESIGN/01-to-be/09-the-node-lifecycle.md @@ -2,7 +2,10 @@ layer: to-be status: designed code: + - mesh-host internal/link/run.go - mesh-host internal/link/enrol.go + - mesh-host packaging/nox-mesh-host-resume.service + - mesh-host packaging/nox-mesh-host-network.sh - mesh-control internal/token - mesh-control internal/inventory/nodes.go updated: 2026-08-27 @@ -650,3 +653,44 @@ the same command against a mesh that is one machine old. - **How a previous declaration is retained and chosen**, which is what rollback of anything else would use ([ADR 0010](../../02-DECISIONS/0010-delivery.md)). - **A node returning after months** applies a very large jump in one go. Correct, and untested. + +## Going away and coming back + +*2026-08-31, from being asked whether a machine that drops off has to be adopted again.* + +**It does not, and nothing about it expires.** A disconnected node is the same node in a different +situation ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) — reachability is state, +not class. The node holds its own identity and the mesh holds the public half; there is no lease, +no timeout, and nothing that lapses while a machine is shut. A laptop closed for a week comes back +and reconnects, and a declaration sent while it was away is waiting for it. + +**The only thing that forces re-enrolment is a machine losing its own key** — a reinstall, an image +re-cloned. That is deliberate: the mesh then reports what it can no longer seal to rather than +delivering blobs the machine cannot open. + +### The gap that was left, and why it mattered + +**A machine that suspends does not know it has been disconnected, and neither does anything else.** +After a resume the socket looks perfectly healthy from inside the process: no error, no close, +because nothing has tried to send anything yet. Heartbeats discover it twenty or thirty seconds +later. + +For those twenty or thirty seconds the node believes it is in the mesh and is not — and *absence +must never be indistinguishable from a failure to answer* is the rule this whole design is built +on. It recovered on its own, which is why this was a quality gap rather than a fault. It was still +the machine waiting to be told something it already knew. + +**So the machine says so.** Waking, and changing network, both rouse the host. + +| | | +|---|---| +| **it ends the current attempt** | not the wait after it — the process is not waiting, it is sitting inside a connection that will never return | +| **by signal, not by anything listening** | a socket for this would be a control surface on every machine, in exchange for saving twenty seconds, and the security argument rests on there not being one | +| **two rouses at once are one** | a machine suspending and resuming repeatedly must not build a backlog of reconnections to work through | +| **the backoff is not reset** | being roused says the machine changed, not that whatever refused the connection has stopped. A laptop woken on a network with no route would otherwise retry at full speed for as long as somebody keeps opening the lid | +| **`down` does not rouse** | the link is already gone, reconnecting will fail, and the backoff exists for exactly that | + +*Checked by holding a link that never returns on its own — which is precisely what a suspended +connection is — rousing it, and requiring the attempt to end and another to begin. And by running +the dispatcher against every event a network manager emits, requiring it to act on the ones that +change where packets go and on no others.*