At 03:30 the store's collector held the registry still while the node-engine's reconcile on that machine was fetching bundle blobs from it: every archive failed 'connection refused' and the machine was held until the next pass. Other machines can meet the same window. A scheduled step now opens its window only once no apply is in flight here (the apply lock is taken just to write the record, so a push still never queues behind the window). An apply whose fetch the store does not answer waits for a window open on its own machine to close, and elsewhere retries with backoff within one bounded budget per apply, well past the window's length; an answer such as 404 still fails at once.
249 lines
9.9 KiB
Go
249 lines
9.9 KiB
Go
package apply
|
|
|
|
// What a maintenance window is holding still, written where every apply on this machine can read it
|
|
// (novox/hq issue 224, ADR 0189).
|
|
//
|
|
// A scheduled step declared `while-stopped` stops its module's containers, runs, and starts them
|
|
// again. Until this, nothing told the **apply** that a window was open, and the apply's rule for a
|
|
// container it finds stopped is the right rule everywhere else: stopped is broken, so remove it and
|
|
// create it running. In the middle of a window that rule reopens the window — for the store's
|
|
// collector, a registry recreated mid-collection accepts an upload the sweep then deletes, and the
|
|
// build that made the image reported success.
|
|
//
|
|
// **The apply learns which containers a window holds; the window does not take the apply lock.**
|
|
// The issue weighed both. Holding the lock for a window makes the race impossible and makes a push
|
|
// that arrives mid-window wait for minutes, which is what issue 185's outages looked like from
|
|
// outside. Telling the apply blocks nothing: a held container is reported as held and left exactly
|
|
// as it is — even when its declaration changed — and the first apply after the window converges it
|
|
// by the ordinary rule. The cost, stated in the issue, is a second source for "is this container
|
|
// meant to be running"; it is kept honest by being short-lived and by expiring on its own.
|
|
//
|
|
// **A file under the node's state directory, not a variable in the scheduler.** The scheduler lives
|
|
// in the daemon, but the daemon is not the only thing that applies here: `mesh-host reconcile` and
|
|
// `mesh-host apply` run as their own processes beside it, and the installer applies the carried
|
|
// bundle — which is where the store's registry comes from — from a third. A map shared in memory
|
|
// would protect only the daemon's own applies, which is to say it would leave open exactly the case
|
|
// a person running `reconcile` by hand at 03:31 creates. The state directory is already where every
|
|
// one of them meets (the apply lock, the node's state, the kept originals), it is root's alone, and
|
|
// each window is one small file written atomically, so no reader ever sees half of one.
|
|
//
|
|
// **A window that is never closed must not hold for ever** — the same risk the field itself carries,
|
|
// one level up. Three things close it:
|
|
//
|
|
// - the scheduler removes the file in a defer that runs after the containers are started again,
|
|
// whatever the step did;
|
|
// - a window whose process is gone is closed: the file names the process that opened it, and a
|
|
// host that crashed or was replaced mid-window is no longer holding anything, so the apply is
|
|
// free to bring back what it left stopped — which is then the apply doing the job the dead
|
|
// process could not;
|
|
// - and every window carries an end, windowAtMost after it opened, past which it is read as
|
|
// closed whoever is still alive. A step that genuinely runs longer than that is pathological, and
|
|
// the apply then does what it did before this record existed, which is the lesser harm next to a
|
|
// service the mesh can never bring back.
|
|
|
|
import (
|
|
"context"
|
|
"encoding/json"
|
|
"errors"
|
|
"fmt"
|
|
"os"
|
|
"path/filepath"
|
|
"sort"
|
|
"strings"
|
|
"syscall"
|
|
"time"
|
|
|
|
"github.com/novox/mesh-host/internal/store"
|
|
)
|
|
|
|
// WindowsName is the directory beside the node's state that holds the open windows, one file each.
|
|
const WindowsName = "windows"
|
|
|
|
// windowAtMost is the longest a window is believed. The store's collector is minutes over a store of
|
|
// fifty-odd repositories; six hours is far past any collection this mesh will do and short enough
|
|
// that a window nothing closed is over by the next working day (novox/hq issue 224).
|
|
const windowAtMost = 6 * time.Hour
|
|
|
|
// Window is one scheduled step holding containers still: which step, which containers by runtime
|
|
// name, since when, and the latest moment it is believed.
|
|
type Window struct {
|
|
Step string `json:"step"`
|
|
Holds []string `json:"holds"`
|
|
Opened time.Time `json:"opened"`
|
|
Until time.Time `json:"until"`
|
|
// PID is the process that opened it. Not reported: it is how a window whose host died is read
|
|
// as closed.
|
|
PID int `json:"pid"`
|
|
}
|
|
|
|
// Windows is where this machine's open windows are recorded. A nil *Windows records nothing and
|
|
// holds nothing — a test, or a caller that has no state directory.
|
|
type Windows struct {
|
|
dir string
|
|
// alive says whether a process is still running. Replaced in tests; the real one asks the kernel.
|
|
alive func(pid int) bool
|
|
}
|
|
|
|
// WindowsIn records windows under stateDir/windows — the directory the node's state lives in, which
|
|
// every process that applies on this machine already shares.
|
|
func WindowsIn(stateDir string) *Windows {
|
|
return &Windows{dir: filepath.Join(stateDir, WindowsName), alive: processAlive}
|
|
}
|
|
|
|
// processAlive is whether pid names a running process. Signal 0 delivers nothing and only asks; a
|
|
// process this one may not signal is still a process (EPERM), which can only happen across users.
|
|
func processAlive(pid int) bool {
|
|
if pid <= 0 {
|
|
return false
|
|
}
|
|
err := syscall.Kill(pid, 0)
|
|
return err == nil || errors.Is(err, syscall.EPERM)
|
|
}
|
|
|
|
// fileFor is where one step's window is written. A step id is the declaration's, which may carry a
|
|
// dot or a slash; the file name keeps what is safe and replaces the rest.
|
|
func (w *Windows) fileFor(step string) string {
|
|
safe := strings.Map(func(r rune) rune {
|
|
switch {
|
|
case r >= 'a' && r <= 'z', r >= 'A' && r <= 'Z', r >= '0' && r <= '9', r == '-', r == '_', r == '.':
|
|
return r
|
|
}
|
|
return '_'
|
|
}, step)
|
|
return filepath.Join(w.dir, "window-"+safe+".json")
|
|
}
|
|
|
|
// open records that step holds these containers still from now, before the first of them is
|
|
// stopped — so no apply can find one stopped and not know why.
|
|
func (w *Windows) open(step string, holds []string, now time.Time) error {
|
|
if w == nil {
|
|
return nil
|
|
}
|
|
if err := os.MkdirAll(w.dir, 0o700); err != nil {
|
|
return err
|
|
}
|
|
raw, err := json.Marshal(Window{Step: step, Holds: holds, Opened: now.UTC(),
|
|
Until: now.UTC().Add(windowAtMost), PID: os.Getpid()})
|
|
if err != nil {
|
|
return err
|
|
}
|
|
return writeAtomically(w.fileFor(step), raw, 0o600)
|
|
}
|
|
|
|
// How long a step about to open a window waits for an apply in flight on this machine to end, and how
|
|
// often it looks (novox/hq issue 291). Past the bound it opens the window anyway, said: the applies
|
|
// then wait for the window instead (store_away.go), which is the slower order, never a failed one.
|
|
var (
|
|
applyWaitAtMost = 15 * time.Minute
|
|
applyPoll = time.Second
|
|
)
|
|
|
|
// openAfterApplies opens step's window only once no apply is in flight on this machine (novox/hq issue
|
|
// 291): an apply that started before the window would otherwise reach the store's archives with its
|
|
// server already held still. The apply lock is held only while the record is written — every apply
|
|
// that starts afterwards reads the window and waits for the store — so a push arriving mid-window
|
|
// still never queues behind the window itself (issue 224's reason for not taking the lock).
|
|
func (w *Windows) openAfterApplies(ctx context.Context, step string, holds []string, log func(string)) error {
|
|
if w == nil {
|
|
return nil
|
|
}
|
|
stateDir := filepath.Dir(w.dir)
|
|
deadline := time.Now().Add(applyWaitAtMost)
|
|
waited := false
|
|
for {
|
|
release, took, err := store.TryLockIn(stateDir)
|
|
if err != nil {
|
|
return err
|
|
}
|
|
if took {
|
|
defer release()
|
|
if waited {
|
|
log(fmt.Sprintf("scheduled step %s: the apply in flight ended; its window opens now", step))
|
|
}
|
|
return w.open(step, holds, time.Now())
|
|
}
|
|
if !waited {
|
|
waited = true
|
|
log(fmt.Sprintf("scheduled step %s: an apply is in flight on this machine — the window opens once "+
|
|
"it ends, so nothing it fetches finds %s held still (novox/hq issue 291)", step, strings.Join(holds, ", ")))
|
|
}
|
|
if time.Now().After(deadline) {
|
|
log(fmt.Sprintf("scheduled step %s: the apply in flight did not end within %s; the window opens "+
|
|
"beside it, and it waits for the window", step, applyWaitAtMost))
|
|
return w.open(step, holds, time.Now())
|
|
}
|
|
select {
|
|
case <-ctx.Done():
|
|
return ctx.Err()
|
|
case <-time.After(applyPoll):
|
|
}
|
|
}
|
|
}
|
|
|
|
// close removes the record of step's window. Called after its containers are started again; a file
|
|
// already gone is the state wanted.
|
|
func (w *Windows) close(step string) error {
|
|
if w == nil {
|
|
return nil
|
|
}
|
|
if err := os.Remove(w.fileFor(step)); err != nil && !errors.Is(err, os.ErrNotExist) {
|
|
return err
|
|
}
|
|
return nil
|
|
}
|
|
|
|
// Open is the windows open at now, by step. One that has ended, or whose process is gone, is not
|
|
// open — and its file is removed on the way, since nothing else will (novox/hq issue 224). A file
|
|
// that cannot be read is not a window either: the alternative is a machine whose containers are
|
|
// never converged because of a byte nobody can see.
|
|
func (w *Windows) Open(now time.Time) []Window {
|
|
if w == nil {
|
|
return nil
|
|
}
|
|
entries, err := os.ReadDir(w.dir)
|
|
if err != nil {
|
|
return nil
|
|
}
|
|
var open []Window
|
|
for _, e := range entries {
|
|
if e.IsDir() || !strings.HasPrefix(e.Name(), "window-") || !strings.HasSuffix(e.Name(), ".json") {
|
|
continue
|
|
}
|
|
path := filepath.Join(w.dir, e.Name())
|
|
raw, err := os.ReadFile(path)
|
|
if err != nil {
|
|
continue
|
|
}
|
|
var win Window
|
|
if err := json.Unmarshal(raw, &win); err != nil {
|
|
_ = os.Remove(path)
|
|
continue
|
|
}
|
|
if !now.Before(win.Until) || !w.alive(win.PID) {
|
|
_ = os.Remove(path)
|
|
continue
|
|
}
|
|
open = append(open, win)
|
|
}
|
|
sort.Slice(open, func(i, j int) bool { return open[i].Step < open[j].Step })
|
|
return open
|
|
}
|
|
|
|
// holding is the open window that holds the container named name, if one does.
|
|
func (w *Windows) holding(name string, now time.Time) (Window, bool) {
|
|
for _, win := range w.Open(now) {
|
|
for _, held := range win.Holds {
|
|
if held == name {
|
|
return win, true
|
|
}
|
|
}
|
|
}
|
|
return Window{}, false
|
|
}
|
|
|
|
// String is the window as a person reads it in a log line.
|
|
func (win Window) String() string {
|
|
return fmt.Sprintf("%s holds %s still since %s", win.Step, strings.Join(win.Holds, ", "),
|
|
win.Opened.Format(time.RFC3339))
|
|
}
|