A service the mesh asked to run is still running a moment later (hq ADR 0184)

The read-back raced the failure: a service manager returns when it has started the process, and a
daemon that refuses its configuration exits a fraction of a second later, so one look saw it alive.
fail2ban took 221ms on the control node and the apply reported "restarted" onto a dead daemon while
both public machines kept no bans at all. The host looks again, after that moment. A unit still
starting is accepted at both looks; a service asked to stop is not waited on.
This commit is contained in:
2026-10-02 18:16:24 +02:00
parent ca7c4a5915
commit 286865dfa7
2 changed files with 128 additions and 2 deletions
+37 -2
View File
@@ -1041,6 +1041,38 @@ type unitReloader interface {
ReloadUnits(ctx context.Context, run system.Runner) error
}
// serviceSettle is how long the host waits before looking at a unit a second time. A test sets it
// to nothing; on a machine it is the window in which a daemon that refuses its configuration dies.
var serviceSettle = 2 * time.Second
// stayedRunning is the state of a unit the host has just asked to run, read twice.
//
// **Because the first read races the failure.** A service manager returns when it has started the
// process, and the unit is "activating" or "active" at that instant whatever the process is about
// to do. A daemon that reads its configuration, refuses it and exits does so a fraction of a second
// later — fail2ban took 221 milliseconds the day this was written — so a single read back says
// running about a machine whose daemon is already gone, and the apply reports "restarted" for a
// service that is dead. Every ban on both public machines was lost that way while every check
// passed (novox/hq ADR 0184), which is the one shape of failure this host exists to refuse.
//
// So it looks again, after the moment in which that happens. It does not wait for a slow unit to
// finish starting: a unit still coming up reads as running both times and is accepted, as before.
// What this catches is a unit that was running and is not any more.
func stayedRunning(ctx context.Context, sys system.System, run Runner, unit string) (string, error) {
state, err := sys.ServiceState(ctx, run, unit)
if err != nil || state != "running" {
return state, err
}
timer := time.NewTimer(serviceSettle)
defer timer.Stop()
select {
case <-ctx.Done():
return state, ctx.Err()
case <-timer.C:
}
return sys.ServiceState(ctx, run, unit)
}
func applyService(ctx context.Context, sys system.System, r *declaration.Service, run Runner,
changed map[string]bool, previous store.Applied) (Outcome, error) {
if r.Stateless() {
@@ -1120,6 +1152,9 @@ func applyService(ctx context.Context, sys system.System, r *declaration.Service
// Read back. A service manager accepting a command says the transaction was accepted,
// not that the unit is running — one that starts and immediately dies satisfies it.
after, err := sys.ServiceState(ctx, run, r.Unit)
if err == nil && r.State == "running" {
after, err = stayedRunning(ctx, sys, run, r.Unit)
}
if err != nil {
return out, err
}
@@ -1140,7 +1175,7 @@ func applyService(ctx context.Context, sys system.System, r *declaration.Service
}
// Read back, for the same reason as above: a unit that starts and immediately dies
// satisfies a service manager and nothing else.
after, err := sys.ServiceState(ctx, run, r.Unit)
after, err := stayedRunning(ctx, sys, run, r.Unit)
if err != nil {
return out, err
}
@@ -1230,7 +1265,7 @@ func reflectOnly(ctx context.Context, sys system.System, r *declaration.Service,
}
// Read back: it was running, and a restart or reload that left it otherwise is a failure —
// the machine's network manager down is not a change to report and move past.
after, err := sys.ServiceState(ctx, run, r.Unit)
after, err := stayedRunning(ctx, sys, run, r.Unit)
if err != nil {
return out, err
}