Review of the ADR 0105 build (hq ADR 0105). The takeover stopped the found
unit and then found out whether the mesh's interface would do; a start that
failed left the machine with no tunnel at all.
Now nothing is stopped until the declared interface listens on the found port
at the found address and the key file it names holds the found key — the
refusal names the remedy — and a mesh interface that fails to start after the
takeover has the found unit started again, with the account saying so. The
account has three states (not taken, taken, down) and is given on every
takeover, failure included. An interface raised by hand is looked at again
for a moment and then refused naming `wg-quick down`. A found unit started
again by hand beside the mesh's is said, not stopped: on the hub it cannot
hold the port, and on a spoke two interfaces with one key would fight.
`mesh-host overlay take --tunnel <iface>` is the path for a node that
enrolled before the mesh knew to take a tunnel over: the found key becomes its
overlay key — identity, sealing and serving keys untouched, so nothing sealed
to the node is remade — and the mesh is told with a rekey signed by the
identity key, over the key left, the key taken and the tunnel. Told first,
written second, so a run again puts right whichever half did not happen.
031: a window of unacknowledged declarations is drained to the newest; the
rest are set aside and reported as superseded. 035: a file resource may say
create-once — written when absent, kept untouched when present (ADR 0087).
054: the bundle installs nftables and loads a base ruleset before the store
and broker, in the table the filter module later replaces (ADR 0088).
A suspended laptop's connection is dead the moment it wakes, and the socket
looks perfectly healthy from inside the process — no error, no close, because
nothing has tried to send anything. Heartbeats find out twenty or thirty
seconds later. For that time the node believes it is in a mesh it has left,
which is the one state this design says must never be indistinguishable from
being connected. The machine knew immediately.
So being roused ends the current attempt rather than only shortening the wait
after it: shortening the wait would do nothing at all, because the process is
not waiting — it is sitting inside a connection that will not return.
A signal, because nothing may listen on a node (novox/hq ADR 0004). A socket
for this would be a control surface on every machine, reachable by anything
that can reach the machine, in exchange for saving twenty seconds — and the
whole security argument rests on there not being one.
Two rouses in the same instant are one: a machine suspending and resuming
repeatedly must not build a backlog of reconnections to work through. And the
backoff is not reset by being roused — that says the machine changed, not that
whatever was refusing the connection has stopped, and a laptop woken on a
network with no route would otherwise retry at full speed for as long as
somebody keeps opening the lid.
The dispatcher acts on the events that change where packets go and not on
`down`: the link is already gone there, reconnecting will fail, and the backoff
exists for exactly that.
The launcher already ran `host run`, and `run` on a machine with no identity
exited with an error. So a freshly installed host, sitting exactly as intended
waiting for somebody to bring it a token, would have counted three failed
starts and rolled back its own installation.
It waits now, and says what it is waiting for. That is the *hosted* state from
the lifecycle: the host is running, it has no identity, and there is nobody to
link to. Every machine passes through it.
An identity that exists and cannot be read is still a fault rather than a wait.
Treating that as "not enrolled yet" would leave a node sitting quietly for ever
while the mesh believes it is a member.
Also: a node now says it is there once a minute. Nothing but its name, because
anything more would be a report, and reports are rare where this is constant --
reading one as the other would make a quiet node look like a stale one. Not
published mandatory, unlike a report: losing one is nothing, the next is a
minute away, and the mesh reads a gap rather than counting arrivals.
Verified in the lab: a node was stopped and the mesh said "out of touch 4m",
then it was started and the mesh said "here" again, without anything else being
touched.
Disconnection is an ordinary situation and not a failure, and until now the
host treated it as the end: the link dropped and the process returned. A laptop
shut for a week would have come back needing somebody to start it again.
Now it reconnects, with a backoff that starts at two seconds and slows to two
minutes. The two common reasons differ in how long they last -- a broker
restarting is back in seconds, a machine that has moved to a network with no
route may be hours -- so it starts fast and slows down, and resets once a
connection has actually held for thirty seconds. Without that reset, a node
that reconnects and immediately drops climbs to the maximum and stays there
long after the cause is gone.
A wrong certificate is said in full every time rather than folded into a retry
count. That does not mean the network is down; it means what answered is not
the mesh this node joined, and no waiting fixes it.
And it says when it gets back in. It logged every failure and nothing on
success, so a log full of "trying again" followed by silence read as still
broken when it meant the opposite.
The other half: a node now keeps what it was told, not only what it applied.
The record of what was applied holds an id, a type and a target -- what removal
needs, not what creation needs -- so it could not be re-applied. The
declaration is kept whole, signed, and verified again every time it is read
back, so the file on disk is trusted for the same reason the message was rather
than for being local. A tampered one is refused, and so is one signed by
another mesh.
With both, the host reconciles against what it was last told every five
minutes, connected or not. That is not polling for changes -- changes are
pushed -- it is the answer to a machine drifting: a file edited by hand, a
container somebody stopped, a service that died.
Verified in the lab. The broker was stopped: the node retried at 2s, 4s, 8s,
saying why each time, and kept its overlay up throughout. The broker came back
and the node rejoined without being touched. A declaration published while a
node was away was waiting on the broker and applied the moment it connected,
which is the buffer ADR 0006 describes doing its job.
04-ISSUES/010. The store now records where each resource came from -- carried,
or declared -- and each origin removes only its own. A declaration removes what
the mesh previously declared and never what the bundle raised.
State written before the field existed reads as carried, because everything a
host had applied by then came from its bundle: there was no other way to tell
it anything. Guessing the other way would have the first upgrade remove the
substrate, which is this fault arriving through the change that fixes it.
Verified on the scenario that caused it, and on the property that had to
survive it: a later declaration dropping a resource still removes that
resource, so removal by omission still means what it meant.
Also stops swallowing a publish failure. A node that applied a declaration and
could not tell the mesh looked exactly like one that had -- the mesh believing
it never answered, the node believing it did, and nothing anywhere saying so.
Reports are published mandatory now, so anything the broker cannot route comes
back and is said out loud rather than dropped in silence.
The loop the whole thing exists for: told, apply, report.
`run` holds one outbound connection open and consumes the node's own queue.
Every declaration is verified against the control plane's signing key before a
byte of it is read as an instruction -- not once at connect, every time. The
transport being pinned is a different question from the instruction being
genuine, and pinning only the first would make the second transitive: a
compromised broker could forge declarations, and this host applies whatever the
link delivers.
Malformed and forged are reported differently, because ADR 0004 requires a host
to tell "this is not from the mesh I joined" from "this is broken". One means
somebody is trying and the other means something needs fixing.
A node now keeps what it needs to come back on its own: the broker's address
and fingerprint, the signing key it believes, and its own broker password --
which the mesh issues at enrolment to replace the token's secret, so the
one-time thing stays one-time and the credential it holds for years is not the
one that was pasted into a terminal.
Verified in the lab end to end. The node enrolled, held its link, received a
signed declaration and applied it -- the file is on the machine with the right
contents, and the host's own record lists both resources.
That run also found issue 010, which is recorded in novox/hq: the declaration
removed every container on the machine, including the control plane that sent
it. Correct reconciliation, shared store, and the first thing that happens.