A node that loses its mesh comes back on its own

Disconnection is an ordinary situation and not a failure, and until now the
host treated it as the end: the link dropped and the process returned. A laptop
shut for a week would have come back needing somebody to start it again.

Now it reconnects, with a backoff that starts at two seconds and slows to two
minutes. The two common reasons differ in how long they last -- a broker
restarting is back in seconds, a machine that has moved to a network with no
route may be hours -- so it starts fast and slows down, and resets once a
connection has actually held for thirty seconds. Without that reset, a node
that reconnects and immediately drops climbs to the maximum and stays there
long after the cause is gone.

A wrong certificate is said in full every time rather than folded into a retry
count. That does not mean the network is down; it means what answered is not
the mesh this node joined, and no waiting fixes it.

And it says when it gets back in. It logged every failure and nothing on
success, so a log full of "trying again" followed by silence read as still
broken when it meant the opposite.

The other half: a node now keeps what it was told, not only what it applied.
The record of what was applied holds an id, a type and a target -- what removal
needs, not what creation needs -- so it could not be re-applied. The
declaration is kept whole, signed, and verified again every time it is read
back, so the file on disk is trusted for the same reason the message was rather
than for being local. A tampered one is refused, and so is one signed by
another mesh.

With both, the host reconciles against what it was last told every five
minutes, connected or not. That is not polling for changes -- changes are
pushed -- it is the answer to a machine drifting: a file edited by hand, a
container somebody stopped, a service that died.

Verified in the lab. The broker was stopped: the node retried at 2s, 4s, 8s,
saying why each time, and kept its overlay up throughout. The broker came back
and the node rejoined without being touched. A declaration published while a
node was away was waiting on the broker and applied the moment it connected,
which is the buffer ADR 0006 describes doing its job.
This commit is contained in:
2026-08-29 20:17:49 +02:00
parent ba31eef80f
commit 980e12a850
5 changed files with 385 additions and 12 deletions
+71 -7
View File
@@ -29,18 +29,77 @@ type Membership struct {
}
// Applier is what the host does with a declaration that has been proved to come from the mesh.
type Applier func(ctx context.Context, declaration []byte) Report
//
// It receives the signature as well as the declaration, so the host can keep both: what it was
// told is kept signed and verified again when it is read back, which means the file on disk is
// trusted for the same reason the message was rather than for being local.
type Applier func(ctx context.Context, declaration, signature []byte) Report
// Announce is how the link says what is happening, so a node running unattended leaves an
// account of it. Nil is allowed and means say nothing.
type Announce func(string)
// Run holds the link open, applying what arrives and reporting what happened.
// Hold keeps this node in the mesh, reconnecting for as long as it is asked to.
//
// Outbound only, and nothing listens on this machine. The connection is the node's presence in
// the mesh: while it is up the node is enrolled, and while it is down the node is disconnected —
// which is an ordinary situation and not a failure, so this returns rather than panicking and
// leaves restarting to whatever supervises it.
// Disconnection is an ordinary situation and not a failure (novox/hq ADR 0004), so this does not
// give up. A laptop shut for a week comes back and reconnects; it does not come back needing
// somebody to start it again.
//
// The backoff exists because the two common reasons differ in how long they last: a broker
// restarting is back in seconds, and a machine that has moved to a network with no route may be
// hours. Retrying every second for hours is a node shouting into nothing; waiting a minute after
// a broker blip is a node that is needlessly late. So it starts fast and slows down, and resets
// once a connection has actually held.
func Hold(ctx context.Context, m Membership, apply Applier, say Announce, timeout time.Duration) error {
const (
first = 2 * time.Second
most = 2 * time.Minute
// A connection that lasted this long counts as having worked, so the next failure starts
// from the bottom again. Without it a node that reconnects and immediately drops climbs
// to the maximum and stays there, long after whatever caused it went away.
settled = 30 * time.Second
)
wait := first
for {
began := time.Now()
err := Run(ctx, m, apply, say, timeout)
if ctx.Err() != nil {
return nil
}
if time.Since(began) > settled {
wait = first
}
switch {
case errors.Is(err, ErrWrongCertificate):
// Said in full every time rather than folded into a retry count. This does not mean
// the network is down; it means what answered is not the mesh this node joined, and
// no amount of waiting fixes it. The node keeps running what it was last told, which
// is the right thing to do while somebody works out what happened.
say("the broker is not the one this node joined: " + err.Error())
say("this will not fix itself. This node keeps running what it was last told.")
case err != nil:
say(fmt.Sprintf("disconnected: %v — trying again in %s", err, wait))
default:
say(fmt.Sprintf("the link closed — trying again in %s", wait))
}
select {
case <-ctx.Done():
return nil
case <-time.After(wait):
}
if wait *= 2; wait > most {
wait = most
}
}
}
// Run holds the link open once, applying what arrives and reporting what happened.
//
// Outbound only, and nothing listens on this machine. Returns when the link ends, for any reason;
// Hold is what decides whether to open it again.
func Run(ctx context.Context, m Membership, apply Applier, say Announce, timeout time.Duration) error {
if say == nil {
say = func(string) {}
@@ -89,6 +148,11 @@ func Run(ctx context.Context, m Membership, apply Applier, say Announce, timeout
if err != nil {
return err
}
// Said, because it is the event anybody watching actually wants. Without it a node logs
// every failure and nothing on success, so a log full of "trying again" and then silence
// reads as still broken when it means the opposite.
say("in the mesh, consuming " + queue)
closed := conn.NotifyClose(make(chan *amqp.Error, 1))
// Published mandatory, so the broker hands back anything it cannot route rather than
@@ -150,7 +214,7 @@ func handleBody(ctx context.Context, m Membership, body []byte, apply Applier) R
if !ed25519.Verify(m.Signer, signed.Declaration, signed.Signature) {
return Report{Node: m.Node, Refused: ErrForged.Error()}
}
return apply(ctx, signed.Declaration)
return apply(ctx, signed.Declaration, signed.Signature)
}
func publishReport(ctx context.Context, channel *amqp.Channel, m Membership, report Report,