The host delivers its own successor, and versions live side by side

The supervision was already right: a clean exit means the host stood aside, and
the launcher's next turn runs what is on disk. Two things made it dead code —
nothing told the running host a successor was waiting, and the rollback resolved
its known-good version through pacman, which no machine here uses and which two
of three operating systems do not have.

Keeping a version rather than a path was the clue. Versions now live in
directories named for them:

- the launcher picks the newest delivered one every time round the loop, or the
  one a rollback pinned, or the host placed by hand when nothing is delivered;
- the running host stands aside between reconciles, never inside one, by exiting
  cleanly — and returns nil so the launcher does not count it as a crash;
- a completed reconcile retires what is older than the predecessor, keeping the
  predecessor because that is what a rollback starts, and never the running one;
- rollback pins the predecessor instead of reinstalling a package: no package
  manager, no cache anyone may clean, same script on every operating system;
- the report says which host version produced it, so 'behind' is answerable.

Newest is when it arrived, never how the name sorts: '1.10' orders before '1.9',
and ordering by name would start an older host and call it an upgrade.

novox/hq ADR 0141. The delivery half — a module carrying the next host — follows;
until then nothing delivers a version and every machine takes the fallback, which
is what it does today.
This commit is contained in:
2026-09-29 00:29:36 +02:00
parent b005d4ef17
commit fbf0fb7d63
8 changed files with 592 additions and 73 deletions
+42 -2
View File
@@ -601,6 +601,20 @@ func runApply(ctx context.Context, opts options, d *declaration.Declaration, raw
" a rollback would have nothing to return to.\n", version, err)
}
}
// And retire what is older than this version's predecessor, on the same evidence known-good is
// written on (novox/hq ADR 0141). The predecessor stays, because it is exactly what a rollback
// starts; everything before it has no reader. Never this version, whatever the answer.
//
// A failure is said and does not fail the apply, for the reason above: what is lost is disk, and
// hiding it would make a machine quietly fill up.
if version != "" {
if retired, err := upgrade.Retire(upgrade.VersionsDir(""), version); err != nil {
fmt.Fprintf(os.Stderr, "mesh-host: applied, and could not retire an older host: %v\n", err)
} else if len(retired) > 0 {
fmt.Fprintf(os.Stderr, "mesh-host: retired the host version(s) %s\n",
strings.Join(retired, ", "))
}
}
// And tell the launcher this start worked. Without it the counter only climbs, and a node
// that has been up for months rolls itself back on its third ordinary restart.
if err := upgrade.ClearAttempts(upgrade.AttemptsPath(opts.state)); err != nil {
@@ -979,12 +993,31 @@ func runLink(ctx context.Context, opts options) error {
sched := apply.NewScheduler(apply.SystemClock(), apply.ExecRunner, say)
go sched.Run(ctx)
// **Standing aside for a successor happens between reconciles and nowhere else** (novox/hq ADR
// 0141). A host that stood aside mid-apply is the half-configured machine this host exists to
// prevent, so the question is asked after an apply has finished and the answer is a clean exit —
// which the launcher already reads as "run whatever is on disk now".
aside, standAside := context.WithCancel(ctx)
defer standAside()
stoodAside := false
applier := func(ctx context.Context, raw, signature []byte) link.Report {
report := applyAndKeep(ctx, opts, raw, &store.Declared{Declaration: raw, Signature: signature}, sched, say)
// **A declaration may carry this machine's membership for another bus.** It arrives as a
// sealed file like any secret, and is read after the rest has applied so the bus it names is
// standing before this machine leaves the one it is on (novox/hq design 28, task 5.2).
adoptDeliveredMembership(identity.Path(opts.state), &mine, say)
switch next, waiting, err := upgrade.Successor(upgrade.VersionsDir(""), version); {
case err != nil:
// Said, not fatal. A host that cannot read the delivered versions is still running this
// machine correctly; what it has lost is the ability to be replaced.
say(fmt.Sprintf("cannot tell whether a newer host is delivered: %v", err))
case waiting:
say(fmt.Sprintf("host %s is delivered; standing aside so the launcher runs it", next.Version))
stoodAside = true
standAside()
}
return report
}
@@ -1015,7 +1048,7 @@ func runLink(ctx context.Context, opts options) error {
}}
})
return link.HoldRoused(ctx, link.Membership{
held := link.HoldRoused(aside, link.Membership{
Node: mine.Node,
Broker: mine.Membership.Broker,
Fingerprint: mine.Membership.Fingerprint,
@@ -1023,6 +1056,13 @@ func runLink(ctx context.Context, opts options) error {
Transport: mine.Membership.Transport,
Signer: mine.Membership.Signer,
}, applier, say, opts.timeout, rousedBySignal(ctx), outbox)
// **Cleanly**, or the launcher counts standing aside as a crash and rolls the new host back
// before it has run once. The context this returns on was cancelled deliberately, so its error
// is not a fault to report.
if stoodAside {
return nil
}
return held
}
// adoptionWatch remembers what the node last said about what it holds and its firewall, so a
@@ -1263,7 +1303,7 @@ func applyAndKeep(ctx context.Context, opts options, raw []byte, signed *store.D
sched.Sync(declared, held)
}
report := link.Report{Carried: carriedPorts(updated), Declared: digestOf(raw)}
report := link.Report{Carried: carriedPorts(updated), Declared: digestOf(raw), Host: version}
// Which of this machine's links face outside, for the filter the mesh writes around them
// (novox/hq ADR 0140). Reported whatever the node's mode: a converged node's filter needs it,
// and an adopted one becomes converged without a further round trip. A machine that cannot read