Files
mesh-host/internal/apply/maintenance_window.go
T
jochen 3c1ac6aef2
mesh/merge-gate pass: builds mesh-host → ace, g14, novox, shanks; no bus step; every machine composes with the change as it did without (4 of 4 compose)
mesh/repo-check pass: its merge-check.sh passed
mesh/delivery delivered
Let a planned maintenance window fail no apply (hq issue 291)
At 03:30 the store's collector held the registry still while the
node-engine's reconcile on that machine was fetching bundle blobs from
it: every archive failed 'connection refused' and the machine was held
until the next pass. Other machines can meet the same window.

A scheduled step now opens its window only once no apply is in flight
here (the apply lock is taken just to write the record, so a push still
never queues behind the window). An apply whose fetch the store does
not answer waits for a window open on its own machine to close, and
elsewhere retries with backoff within one bounded budget per apply,
well past the window's length; an answer such as 404 still fails at
once.
2026-10-07 13:11:03 +02:00

249 lines
9.9 KiB
Go

package apply
// What a maintenance window is holding still, written where every apply on this machine can read it
// (novox/hq issue 224, ADR 0189).
//
// A scheduled step declared `while-stopped` stops its module's containers, runs, and starts them
// again. Until this, nothing told the **apply** that a window was open, and the apply's rule for a
// container it finds stopped is the right rule everywhere else: stopped is broken, so remove it and
// create it running. In the middle of a window that rule reopens the window — for the store's
// collector, a registry recreated mid-collection accepts an upload the sweep then deletes, and the
// build that made the image reported success.
//
// **The apply learns which containers a window holds; the window does not take the apply lock.**
// The issue weighed both. Holding the lock for a window makes the race impossible and makes a push
// that arrives mid-window wait for minutes, which is what issue 185's outages looked like from
// outside. Telling the apply blocks nothing: a held container is reported as held and left exactly
// as it is — even when its declaration changed — and the first apply after the window converges it
// by the ordinary rule. The cost, stated in the issue, is a second source for "is this container
// meant to be running"; it is kept honest by being short-lived and by expiring on its own.
//
// **A file under the node's state directory, not a variable in the scheduler.** The scheduler lives
// in the daemon, but the daemon is not the only thing that applies here: `mesh-host reconcile` and
// `mesh-host apply` run as their own processes beside it, and the installer applies the carried
// bundle — which is where the store's registry comes from — from a third. A map shared in memory
// would protect only the daemon's own applies, which is to say it would leave open exactly the case
// a person running `reconcile` by hand at 03:31 creates. The state directory is already where every
// one of them meets (the apply lock, the node's state, the kept originals), it is root's alone, and
// each window is one small file written atomically, so no reader ever sees half of one.
//
// **A window that is never closed must not hold for ever** — the same risk the field itself carries,
// one level up. Three things close it:
//
// - the scheduler removes the file in a defer that runs after the containers are started again,
// whatever the step did;
// - a window whose process is gone is closed: the file names the process that opened it, and a
// host that crashed or was replaced mid-window is no longer holding anything, so the apply is
// free to bring back what it left stopped — which is then the apply doing the job the dead
// process could not;
// - and every window carries an end, windowAtMost after it opened, past which it is read as
// closed whoever is still alive. A step that genuinely runs longer than that is pathological, and
// the apply then does what it did before this record existed, which is the lesser harm next to a
// service the mesh can never bring back.
import (
"context"
"encoding/json"
"errors"
"fmt"
"os"
"path/filepath"
"sort"
"strings"
"syscall"
"time"
"github.com/novox/mesh-host/internal/store"
)
// WindowsName is the directory beside the node's state that holds the open windows, one file each.
const WindowsName = "windows"
// windowAtMost is the longest a window is believed. The store's collector is minutes over a store of
// fifty-odd repositories; six hours is far past any collection this mesh will do and short enough
// that a window nothing closed is over by the next working day (novox/hq issue 224).
const windowAtMost = 6 * time.Hour
// Window is one scheduled step holding containers still: which step, which containers by runtime
// name, since when, and the latest moment it is believed.
type Window struct {
Step string `json:"step"`
Holds []string `json:"holds"`
Opened time.Time `json:"opened"`
Until time.Time `json:"until"`
// PID is the process that opened it. Not reported: it is how a window whose host died is read
// as closed.
PID int `json:"pid"`
}
// Windows is where this machine's open windows are recorded. A nil *Windows records nothing and
// holds nothing — a test, or a caller that has no state directory.
type Windows struct {
dir string
// alive says whether a process is still running. Replaced in tests; the real one asks the kernel.
alive func(pid int) bool
}
// WindowsIn records windows under stateDir/windows — the directory the node's state lives in, which
// every process that applies on this machine already shares.
func WindowsIn(stateDir string) *Windows {
return &Windows{dir: filepath.Join(stateDir, WindowsName), alive: processAlive}
}
// processAlive is whether pid names a running process. Signal 0 delivers nothing and only asks; a
// process this one may not signal is still a process (EPERM), which can only happen across users.
func processAlive(pid int) bool {
if pid <= 0 {
return false
}
err := syscall.Kill(pid, 0)
return err == nil || errors.Is(err, syscall.EPERM)
}
// fileFor is where one step's window is written. A step id is the declaration's, which may carry a
// dot or a slash; the file name keeps what is safe and replaces the rest.
func (w *Windows) fileFor(step string) string {
safe := strings.Map(func(r rune) rune {
switch {
case r >= 'a' && r <= 'z', r >= 'A' && r <= 'Z', r >= '0' && r <= '9', r == '-', r == '_', r == '.':
return r
}
return '_'
}, step)
return filepath.Join(w.dir, "window-"+safe+".json")
}
// open records that step holds these containers still from now, before the first of them is
// stopped — so no apply can find one stopped and not know why.
func (w *Windows) open(step string, holds []string, now time.Time) error {
if w == nil {
return nil
}
if err := os.MkdirAll(w.dir, 0o700); err != nil {
return err
}
raw, err := json.Marshal(Window{Step: step, Holds: holds, Opened: now.UTC(),
Until: now.UTC().Add(windowAtMost), PID: os.Getpid()})
if err != nil {
return err
}
return writeAtomically(w.fileFor(step), raw, 0o600)
}
// How long a step about to open a window waits for an apply in flight on this machine to end, and how
// often it looks (novox/hq issue 291). Past the bound it opens the window anyway, said: the applies
// then wait for the window instead (store_away.go), which is the slower order, never a failed one.
var (
applyWaitAtMost = 15 * time.Minute
applyPoll = time.Second
)
// openAfterApplies opens step's window only once no apply is in flight on this machine (novox/hq issue
// 291): an apply that started before the window would otherwise reach the store's archives with its
// server already held still. The apply lock is held only while the record is written — every apply
// that starts afterwards reads the window and waits for the store — so a push arriving mid-window
// still never queues behind the window itself (issue 224's reason for not taking the lock).
func (w *Windows) openAfterApplies(ctx context.Context, step string, holds []string, log func(string)) error {
if w == nil {
return nil
}
stateDir := filepath.Dir(w.dir)
deadline := time.Now().Add(applyWaitAtMost)
waited := false
for {
release, took, err := store.TryLockIn(stateDir)
if err != nil {
return err
}
if took {
defer release()
if waited {
log(fmt.Sprintf("scheduled step %s: the apply in flight ended; its window opens now", step))
}
return w.open(step, holds, time.Now())
}
if !waited {
waited = true
log(fmt.Sprintf("scheduled step %s: an apply is in flight on this machine — the window opens once "+
"it ends, so nothing it fetches finds %s held still (novox/hq issue 291)", step, strings.Join(holds, ", ")))
}
if time.Now().After(deadline) {
log(fmt.Sprintf("scheduled step %s: the apply in flight did not end within %s; the window opens "+
"beside it, and it waits for the window", step, applyWaitAtMost))
return w.open(step, holds, time.Now())
}
select {
case <-ctx.Done():
return ctx.Err()
case <-time.After(applyPoll):
}
}
}
// close removes the record of step's window. Called after its containers are started again; a file
// already gone is the state wanted.
func (w *Windows) close(step string) error {
if w == nil {
return nil
}
if err := os.Remove(w.fileFor(step)); err != nil && !errors.Is(err, os.ErrNotExist) {
return err
}
return nil
}
// Open is the windows open at now, by step. One that has ended, or whose process is gone, is not
// open — and its file is removed on the way, since nothing else will (novox/hq issue 224). A file
// that cannot be read is not a window either: the alternative is a machine whose containers are
// never converged because of a byte nobody can see.
func (w *Windows) Open(now time.Time) []Window {
if w == nil {
return nil
}
entries, err := os.ReadDir(w.dir)
if err != nil {
return nil
}
var open []Window
for _, e := range entries {
if e.IsDir() || !strings.HasPrefix(e.Name(), "window-") || !strings.HasSuffix(e.Name(), ".json") {
continue
}
path := filepath.Join(w.dir, e.Name())
raw, err := os.ReadFile(path)
if err != nil {
continue
}
var win Window
if err := json.Unmarshal(raw, &win); err != nil {
_ = os.Remove(path)
continue
}
if !now.Before(win.Until) || !w.alive(win.PID) {
_ = os.Remove(path)
continue
}
open = append(open, win)
}
sort.Slice(open, func(i, j int) bool { return open[i].Step < open[j].Step })
return open
}
// holding is the open window that holds the container named name, if one does.
func (w *Windows) holding(name string, now time.Time) (Window, bool) {
for _, win := range w.Open(now) {
for _, held := range win.Holds {
if held == name {
return win, true
}
}
}
return Window{}, false
}
// String is the window as a person reads it in a log line.
func (win Window) String() string {
return fmt.Sprintf("%s holds %s still since %s", win.Step, strings.Join(win.Holds, ", "),
win.Opened.Format(time.RFC3339))
}