apply: a scheduled step is a container run on a cadence (ADR 0053)
The recurring twin of run-once, one modifier over: a container marked schedule: "<cron>" is run to completion on its cadence, not started as a service and not run once as a gate. The gating rule is deliberately reversed. Installing a schedule records it as present state and reports the node current at once (applySchedule) -- it never runs the container and does not gate what follows. A Scheduler, held for the life of the daemon and re-established from each applied declaration (the declaration is the source of truth, ADR 0018), fires the container off an injected clock. A run that exits non-zero is logged and never fails the apply or flips the node's state, because it happens outside the apply and the store entirely. Runs never stack: a run still going when the next is due is skipped, not started as a second copy. No new host shape and no new action -- schedule is a string on the container the host already has, and the host process runs the container itself rather than installing a system timer (the rejected option 1). A minimal five-field cron (declaration/cron.go) validates on arrival and computes the next due minute; time is injected so the scheduler is tested without the wall clock. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
This commit is contained in:
+71
-12
@@ -874,6 +874,12 @@ func containerSpec(r *declaration.Container) string {
|
||||
for _, a := range r.Args {
|
||||
b.WriteString("arg " + a + "\n")
|
||||
}
|
||||
// The cadence is part of what was declared, so a changed schedule is a changed spec — the marker
|
||||
// moves and the install is reported "updated" and re-established. Added only when present, so no
|
||||
// ordinary container's or run-once step's digest moves for a field it does not set.
|
||||
if r.Schedule != "" {
|
||||
b.WriteString("schedule " + r.Schedule + "\n")
|
||||
}
|
||||
return fmt.Sprintf("%x", sha256.Sum256([]byte(b.String())))
|
||||
}
|
||||
|
||||
@@ -942,6 +948,16 @@ func applyContainer(ctx context.Context, r *declaration.Container, run Runner, c
|
||||
out := begin(r)
|
||||
want := containerSpec(r)
|
||||
|
||||
// A scheduled step is state that is present, not a container to start (novox/hq ADR 0053).
|
||||
// Installing it records the schedule and reports the node current at once — the deliberate
|
||||
// inversion of run-once, which gates. The recurring run is fired by the host's Scheduler off the
|
||||
// clock, re-established from this declaration each apply, and NEVER here — so installing does not
|
||||
// run the container and does not even need a runtime present. Checked before the runtime probe
|
||||
// for exactly that reason.
|
||||
if r.Schedule != "" {
|
||||
return applySchedule(r, want, previous)
|
||||
}
|
||||
|
||||
cri, err := containerRuntime(ctx, run)
|
||||
if err != nil {
|
||||
return out, fmt.Errorf("%w, so nothing can be said about %q", err, r.Name)
|
||||
@@ -1069,6 +1085,29 @@ func applyRunOnce(ctx context.Context, r *declaration.Container, run Runner, cri
|
||||
|
||||
// Run in the foreground so the runtime waits for the container and hands back its exit code.
|
||||
// No --detach and no --restart: a step that is restarted is not a step.
|
||||
args := foregroundRunArgs(r, want)
|
||||
|
||||
if _, err := run(ctx, cri, args...); err != nil {
|
||||
// A non-zero exit or a runtime that could not start it. Either way the step did not make
|
||||
// the machine ready, so the apply must not go on to the container that needs it.
|
||||
return out, fmt.Errorf("run-once step %s did not complete: %w", r.Name, err)
|
||||
}
|
||||
|
||||
// It completed. Remove the exited container so a later apply is not confused by a stopped one;
|
||||
// the record that it ran is the digest below, which the caller persists after the fact.
|
||||
_, _ = run(ctx, cri, "rm", "-f", r.Name)
|
||||
|
||||
out.Action = "created"
|
||||
out.Detail = "run-once step completed"
|
||||
out.wrote = want
|
||||
return out, nil
|
||||
}
|
||||
|
||||
// foregroundRunArgs builds a `docker run` that runs a container to completion and hands back its
|
||||
// exit code — no --detach, no --restart, because a step that is restarted is not a step. Shared by
|
||||
// a run-once step (novox/hq ADR 0052) and by one fire of a scheduled step (ADR 0053), which are the
|
||||
// same "run the container and let it exit" up to how often it happens.
|
||||
func foregroundRunArgs(r *declaration.Container, want string) []string {
|
||||
args := []string{"run", "--name", r.Name}
|
||||
for _, file := range r.EnvFile {
|
||||
args = append(args, "--env-file", file)
|
||||
@@ -1088,20 +1127,40 @@ func applyRunOnce(ctx context.Context, r *declaration.Container, run Runner, cri
|
||||
}
|
||||
args = append(args, r.Image)
|
||||
args = append(args, r.Args...)
|
||||
return args
|
||||
}
|
||||
|
||||
if _, err := run(ctx, cri, args...); err != nil {
|
||||
// A non-zero exit or a runtime that could not start it. Either way the step did not make
|
||||
// the machine ready, so the apply must not go on to the container that needs it.
|
||||
return out, fmt.Errorf("run-once step %s did not complete: %w", r.Name, err)
|
||||
}
|
||||
|
||||
// It completed. Remove the exited container so a later apply is not confused by a stopped one;
|
||||
// the record that it ran is the digest below, which the caller persists after the fact.
|
||||
_, _ = run(ctx, cri, "rm", "-f", r.Name)
|
||||
|
||||
out.Action = "created"
|
||||
out.Detail = "run-once step completed"
|
||||
// applySchedule installs a scheduled step: it records the schedule as present and reports the node
|
||||
// current, without running anything (novox/hq ADR 0053).
|
||||
//
|
||||
// This is the deliberate inversion of run-once. A run-once step gates the apply — it runs to
|
||||
// completion here and a non-zero exit halts everything after it — because "seed the store before the
|
||||
// broker starts" is a precondition of convergence. A scheduled step is the opposite: it runs *after*
|
||||
// the machine is up, on its own clock, and a single failed run is an ordinary operational event. So
|
||||
// installing it is pure state: the schedule is present, like a running service, and the apply is
|
||||
// current at once. The recurring run is fired by the host's Scheduler off the clock (see
|
||||
// schedule.go), re-established from the applied declaration each pass because the declaration is the
|
||||
// source of truth (ADR 0018) — never from here, and never persisted beyond what the mesh already
|
||||
// owns.
|
||||
//
|
||||
// The marker is the declaration's digest, exactly as for run-once, so a re-apply of the same
|
||||
// declaration reports the schedule unchanged and a changed image, environment or cadence reports it
|
||||
// re-installed. A failed run touches none of this: it happens entirely in the Scheduler, outside the
|
||||
// apply and the store, which is why a routine job's failure can never flip the node's state.
|
||||
func applySchedule(r *declaration.Container, want string, previous store.Applied) (Outcome, error) {
|
||||
out := begin(r)
|
||||
out.wrote = want
|
||||
switch {
|
||||
case previous.Wrote == want:
|
||||
out.Action = "unchanged"
|
||||
out.Detail = "scheduled step; already installed for this declaration"
|
||||
case previous.Wrote != "":
|
||||
out.Action = "updated"
|
||||
out.Detail = "scheduled step re-installed; its image, environment or cadence changed"
|
||||
default:
|
||||
out.Action = "created"
|
||||
out.Detail = "scheduled step installed; the host runs it on its cadence"
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user