apply: a scheduled step is a container run on a cadence (ADR 0053)

The recurring twin of run-once, one modifier over: a container marked
schedule: "<cron>" is run to completion on its cadence, not started as a
service and not run once as a gate.

The gating rule is deliberately reversed. Installing a schedule records it
as present state and reports the node current at once (applySchedule) --
it never runs the container and does not gate what follows. A Scheduler,
held for the life of the daemon and re-established from each applied
declaration (the declaration is the source of truth, ADR 0018), fires the
container off an injected clock. A run that exits non-zero is logged and
never fails the apply or flips the node's state, because it happens
outside the apply and the store entirely. Runs never stack: a run still
going when the next is due is skipped, not started as a second copy.

No new host shape and no new action -- schedule is a string on the
container the host already has, and the host process runs the container
itself rather than installing a system timer (the rejected option 1). A
minimal five-field cron (declaration/cron.go) validates on arrival and
computes the next due minute; time is injected so the scheduler is tested
without the wall clock.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
This commit is contained in:
2026-09-06 14:08:44 +02:00
parent 24e9ae4065
commit 9d1f001dcc
8 changed files with 1076 additions and 22 deletions
+71 -12
View File
@@ -874,6 +874,12 @@ func containerSpec(r *declaration.Container) string {
for _, a := range r.Args {
b.WriteString("arg " + a + "\n")
}
// The cadence is part of what was declared, so a changed schedule is a changed spec — the marker
// moves and the install is reported "updated" and re-established. Added only when present, so no
// ordinary container's or run-once step's digest moves for a field it does not set.
if r.Schedule != "" {
b.WriteString("schedule " + r.Schedule + "\n")
}
return fmt.Sprintf("%x", sha256.Sum256([]byte(b.String())))
}
@@ -942,6 +948,16 @@ func applyContainer(ctx context.Context, r *declaration.Container, run Runner, c
out := begin(r)
want := containerSpec(r)
// A scheduled step is state that is present, not a container to start (novox/hq ADR 0053).
// Installing it records the schedule and reports the node current at once — the deliberate
// inversion of run-once, which gates. The recurring run is fired by the host's Scheduler off the
// clock, re-established from this declaration each apply, and NEVER here — so installing does not
// run the container and does not even need a runtime present. Checked before the runtime probe
// for exactly that reason.
if r.Schedule != "" {
return applySchedule(r, want, previous)
}
cri, err := containerRuntime(ctx, run)
if err != nil {
return out, fmt.Errorf("%w, so nothing can be said about %q", err, r.Name)
@@ -1069,6 +1085,29 @@ func applyRunOnce(ctx context.Context, r *declaration.Container, run Runner, cri
// Run in the foreground so the runtime waits for the container and hands back its exit code.
// No --detach and no --restart: a step that is restarted is not a step.
args := foregroundRunArgs(r, want)
if _, err := run(ctx, cri, args...); err != nil {
// A non-zero exit or a runtime that could not start it. Either way the step did not make
// the machine ready, so the apply must not go on to the container that needs it.
return out, fmt.Errorf("run-once step %s did not complete: %w", r.Name, err)
}
// It completed. Remove the exited container so a later apply is not confused by a stopped one;
// the record that it ran is the digest below, which the caller persists after the fact.
_, _ = run(ctx, cri, "rm", "-f", r.Name)
out.Action = "created"
out.Detail = "run-once step completed"
out.wrote = want
return out, nil
}
// foregroundRunArgs builds a `docker run` that runs a container to completion and hands back its
// exit code — no --detach, no --restart, because a step that is restarted is not a step. Shared by
// a run-once step (novox/hq ADR 0052) and by one fire of a scheduled step (ADR 0053), which are the
// same "run the container and let it exit" up to how often it happens.
func foregroundRunArgs(r *declaration.Container, want string) []string {
args := []string{"run", "--name", r.Name}
for _, file := range r.EnvFile {
args = append(args, "--env-file", file)
@@ -1088,20 +1127,40 @@ func applyRunOnce(ctx context.Context, r *declaration.Container, run Runner, cri
}
args = append(args, r.Image)
args = append(args, r.Args...)
return args
}
if _, err := run(ctx, cri, args...); err != nil {
// A non-zero exit or a runtime that could not start it. Either way the step did not make
// the machine ready, so the apply must not go on to the container that needs it.
return out, fmt.Errorf("run-once step %s did not complete: %w", r.Name, err)
}
// It completed. Remove the exited container so a later apply is not confused by a stopped one;
// the record that it ran is the digest below, which the caller persists after the fact.
_, _ = run(ctx, cri, "rm", "-f", r.Name)
out.Action = "created"
out.Detail = "run-once step completed"
// applySchedule installs a scheduled step: it records the schedule as present and reports the node
// current, without running anything (novox/hq ADR 0053).
//
// This is the deliberate inversion of run-once. A run-once step gates the apply — it runs to
// completion here and a non-zero exit halts everything after it — because "seed the store before the
// broker starts" is a precondition of convergence. A scheduled step is the opposite: it runs *after*
// the machine is up, on its own clock, and a single failed run is an ordinary operational event. So
// installing it is pure state: the schedule is present, like a running service, and the apply is
// current at once. The recurring run is fired by the host's Scheduler off the clock (see
// schedule.go), re-established from the applied declaration each pass because the declaration is the
// source of truth (ADR 0018) — never from here, and never persisted beyond what the mesh already
// owns.
//
// The marker is the declaration's digest, exactly as for run-once, so a re-apply of the same
// declaration reports the schedule unchanged and a changed image, environment or cadence reports it
// re-installed. A failed run touches none of this: it happens entirely in the Scheduler, outside the
// apply and the store, which is why a routine job's failure can never flip the node's state.
func applySchedule(r *declaration.Container, want string, previous store.Applied) (Outcome, error) {
out := begin(r)
out.wrote = want
switch {
case previous.Wrote == want:
out.Action = "unchanged"
out.Detail = "scheduled step; already installed for this declaration"
case previous.Wrote != "":
out.Action = "updated"
out.Detail = "scheduled step re-installed; its image, environment or cadence changed"
default:
out.Action = "created"
out.Detail = "scheduled step installed; the host runs it on its cadence"
}
return out, nil
}