A machine that wakes or moves says so, instead of waiting to be told

A suspended laptop's connection is dead the moment it wakes, and the socket
looks perfectly healthy from inside the process — no error, no close, because
nothing has tried to send anything. Heartbeats find out twenty or thirty
seconds later. For that time the node believes it is in a mesh it has left,
which is the one state this design says must never be indistinguishable from
being connected. The machine knew immediately.

So being roused ends the current attempt rather than only shortening the wait
after it: shortening the wait would do nothing at all, because the process is
not waiting — it is sitting inside a connection that will not return.

A signal, because nothing may listen on a node (novox/hq ADR 0004). A socket
for this would be a control surface on every machine, reachable by anything
that can reach the machine, in exchange for saving twenty seconds — and the
whole security argument rests on there not being one.

Two rouses in the same instant are one: a machine suspending and resuming
repeatedly must not build a backlog of reconnections to work through. And the
backoff is not reset by being roused — that says the machine changed, not that
whatever was refusing the connection has stopped, and a laptop woken on a
network with no route would otherwise retry at full speed for as long as
somebody keeps opening the lid.

The dispatcher acts on the events that change where packets go and not on
`down`: the link is already gone there, reconnecting will fail, and the backoff
exists for exactly that.
This commit is contained in:
2026-08-31 10:21:26 +02:00
parent 5bc0006e83
commit 8fcfa88fe0
8 changed files with 302 additions and 3 deletions
+22
View File
@@ -0,0 +1,22 @@
#!/bin/sh
# Rouse the host when this machine's network changes.
#
# Installed as a NetworkManager dispatcher script (/etc/NetworkManager/dispatcher.d) and as a
# networkd-dispatcher one. Both hand the interface and the event as arguments; both are ignored
# beyond the event, because *which* interface changed does not matter — what matters is that a
# connection opened over the old route is now pointing at nothing, and that is true whichever
# interface it was.
#
# A machine that moves from wifi to ethernet holds a socket that looks perfectly healthy from
# inside the process: no error, no close, because nothing has tried to send anything. Heartbeats
# find it twenty or thirty seconds later. The machine knew immediately.
set -eu
event="${2:-}"
case "$event" in
up|dhcp4-change|dhcp6-change|connectivity-change|routes-change)
# Only events that can change where packets go. `down` is deliberately not one: the link is
# already gone, reconnecting will fail, and the backoff exists for exactly that.
systemctl start --no-block nox-mesh-host-roused.service 2>/dev/null || true
;;
esac
+14
View File
@@ -0,0 +1,14 @@
# Rouse the host when this machine wakes.
#
# After sleep.target rather than before: the point is to act once the machine is back, and a
# signal sent on the way down would be read by a process that is about to be frozen with it.
[Unit]
Description=Rouse the Novox Mesh host after resume
After=suspend.target hibernate.target hybrid-sleep.target suspend-then-hibernate.target
[Service]
Type=oneshot
ExecStart=/bin/sh -c 'systemctl start --no-block nox-mesh-host-roused.service'
[Install]
WantedBy=suspend.target hibernate.target hybrid-sleep.target suspend-then-hibernate.target
+17
View File
@@ -0,0 +1,17 @@
# Tell the running host that this machine's link is probably stale.
#
# **A laptop knows it just woke; the link does not.** After a resume the socket looks perfectly
# healthy from inside the process — no error, no close, because nothing has tried to send
# anything. Heartbeats discover it twenty or thirty seconds later, and for that time the node
# believes it is in a mesh it has left.
#
# A signal rather than anything that listens: nothing may listen on a node (novox/hq ADR 0004),
# and a socket for this would be a control surface on every machine in exchange for saving twenty
# seconds.
[Unit]
Description=Tell the Novox Mesh host its link may be stale
[Service]
Type=oneshot
# Nothing to do if the host is not running: this is a hint to a process, not a way to start one.
ExecStart=/bin/sh -c 'systemctl is-active --quiet nox-mesh-host.service && systemctl kill --signal=SIGHUP nox-mesh-host.service || true'
+42
View File
@@ -0,0 +1,42 @@
#!/bin/sh
# The dispatcher script rouses on the events that change where packets go, and on nothing else.
#
# Checked by running it with a stub systemctl on PATH, because the thing worth testing is which
# events it acts on — a script that rouses on `down` would reconnect into a network that is gone,
# and one that rouses on nothing is the timeout it was written to avoid.
set -eu
here=$(cd "$(dirname "$0")" && pwd)
work=$(mktemp -d)
trap 'rm -rf "$work"' EXIT
mkdir -p "$work/bin"
cat > "$work/bin/systemctl" <<'STUB'
#!/bin/sh
echo "$@" >> "$ROUSED_LOG"
STUB
chmod +x "$work/bin/systemctl"
export PATH="$work/bin:$PATH"
export ROUSED_LOG="$work/rousings"
: > "$ROUSED_LOG"
for event in up dhcp4-change connectivity-change routes-change; do
sh "$here/nox-mesh-host-network.sh" wlan0 "$event"
done
count=$(grep -c "nox-mesh-host-roused" "$ROUSED_LOG" || true)
if [ "$count" -ne 4 ]; then
echo "FAIL: four events that change where packets go roused $count time(s)" >&2
exit 1
fi
: > "$ROUSED_LOG"
for event in down pre-up hostname ""; do
sh "$here/nox-mesh-host-network.sh" wlan0 "$event"
done
if [ -s "$ROUSED_LOG" ]; then
echo "FAIL: an event that does not change where packets go roused the host:" >&2
cat "$ROUSED_LOG" >&2
exit 1
fi
echo "the dispatcher rouses on route changes and nothing else"