The recurring twin of run-once (0052): a schedule modifier on the container shape, reusing its security bound (no new host action/shape, strictly less than an action) and reversing its gating rule — a scheduled step runs after convergence, does not gate the apply, and a failed run is logged, not fatal. Unblocks kometa's sync and pollers. Accepted per direction to build the primitive now. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
182 lines
13 KiB
Markdown
182 lines
13 KiB
Markdown
---
|
|
topic: what runs on it
|
|
status: accepted
|
|
date: 2026-09-06
|
|
deciders: jochen
|
|
reconstructed: false
|
|
extends: 0052-a-step-that-runs-once-before-a-container.md
|
|
---
|
|
|
|
# 53. A scheduled step is a container run on a recurring schedule
|
|
|
|
## Context
|
|
|
|
**[ADR 0052](0052-a-step-that-runs-once-before-a-container.md) gave the mesh a step that runs *once*;
|
|
a real class of modules needs one that runs *again and again*.** kometa reconciles a media library
|
|
against its lists on a timer; a ticketing integration polls its source for new work every few minutes;
|
|
a backup, a cache warm, a metrics roll-up all recur. The host's vocabulary describes *state* — a file,
|
|
a directory, a container that should be running — and 0052 added *a step that happens once and is
|
|
done*. Neither says *this should happen every night at 3, forever*
|
|
([04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md) named
|
|
the lifecycle gap; 0052 closed the once-before-start face of it and left the recurring face open).
|
|
|
|
**The modules that need it already run their own code.** As with run-once, the ground has shifted
|
|
since the old mesh's flaky hooks: a module with tools or events runs a **process of its own** — a
|
|
container carrying the module's compiled code under the single scoped account the mesh gave it
|
|
([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The work that
|
|
must recur is already in that image, under that account. The mesh does not need a way to run a
|
|
module's code on a timer — it has the code and the account. It needs a way to say **run this
|
|
container again on this cadence**.
|
|
|
|
**The shape of the answer is already decided, one modifier over.** 0052 rejected a host `run` command,
|
|
per-phase hooks, and a distinct one-shot resource type, and adopted *a modifier on the `container`
|
|
shape the host already has*, because a one-shot is a container that happens to exit. A scheduled step
|
|
is the same container that happens to exit — run repeatedly. Reusing that shape keeps the host
|
|
vocabulary flat and inherits 0052's security bound whole. The only thing 0052's `run-once` does not
|
|
carry is *when to run it again*.
|
|
|
|
**One difference from run-once changes a rule, and it is the reason this is its own record.** A
|
|
run-once step **gates the apply**: it is placed before the container that depends on it, and a failure
|
|
halts everything after it, because "seed the store before the broker starts" is a correctness
|
|
precondition ([ADR 0052](0052-a-step-that-runs-once-before-a-container.md)). A scheduled step is the
|
|
opposite: it runs *after* the machine is up and converged, on its own clock, and a single failed run
|
|
is an ordinary operational event — the next run comes anyway. A scheduled step that halted the apply,
|
|
or that a failed run marked the node not-current over, would make a routine poll into a reason the
|
|
whole machine reads as broken. So the gating rule 0052 established is exactly the rule this record
|
|
must **not** inherit.
|
|
|
|
## Considered Options
|
|
|
|
1. **A host `cron`/`timer` shape — the host installs a system timer that runs a command.** The direct
|
|
reading. **Rejected**, for the reason 0052 rejected a host `run` shape: it widens what a
|
|
*compromised control plane* can express toward "run this command on the machine, forever," which is
|
|
the largest widening there is, and it is `action`-from-the-link by another name
|
|
([ADR 0005](0005-the-node-host.md)). A recurring command is worse than a one-off, because it
|
|
persists.
|
|
|
|
2. **Per-phase lifecycle hooks** — `on-schedule` joining `pre-start`/`post-start` as named code the
|
|
mesh runs. **Rejected for now**, as in 0052: it is the flaky hook engine the issue warns against,
|
|
and the three modules that need this need one thing — *run this container on a cadence* — which a
|
|
narrow modifier expresses without deciding a whole hook vocabulary.
|
|
|
|
3. **A distinct `scheduled` resource type**, sibling to `container`. **Rejected**, as 0052 rejected a
|
|
distinct one-shot type: it spends a whole new host shape (a `Type`, a struct, an applier, a place
|
|
in every host's shape list) on something a `container` already almost is — a scheduled task *is* a
|
|
container (pinned image, account, volumes, environment) that runs on a clock.
|
|
|
|
4. **A modifier on the existing `container` shape: `schedule`, a cron expression the host runs the
|
|
container on.** **Adopted.** It reuses the shape the host has, adds no new host action, and sits
|
|
beside `run-once` as its recurring twin — the same container, exited, run again.
|
|
|
|
## Decision
|
|
|
|
**A scheduled step is an ordinary `container`, marked with a `schedule`.** The manifest sets
|
|
`schedule: "<cron>"` on a container resource — a standard five-field cron expression. Everything else
|
|
about it is a container as before: a digest-pinned image
|
|
([ADR 0006](0006-the-substrate-and-the-control-plane.md)), volumes, an environment, and for a module's
|
|
own code the same scoped account its runtime already holds
|
|
([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The mesh adds no
|
|
way to run code; it marks a container the host must **run on that cadence, each time to completion**,
|
|
rather than start once and leave running (a service) or run once and gate (a run-once step).
|
|
|
|
**A run is fired by the clock, not by the apply, and does not gate it.** Applying the declaration
|
|
installs the schedule; it does not run the step. The machine converges — reports applied and current —
|
|
as soon as the schedule is installed, exactly as it does for a service that is running. Thereafter the
|
|
host fires the container when the cron expression is due. This is the deliberate inversion of 0052:
|
|
a scheduled step is downstream of convergence, not a precondition of it.
|
|
|
|
**A failed run is recorded and the next run still comes; it never marks the node not-current.** A run
|
|
that exits non-zero is logged against the module — visible, auditable — but it does not fail the apply,
|
|
does not halt other resources, and does not flip the node's reported state. A poll that fails at 03:00
|
|
and succeeds at 03:05 is the system working, not a machine that needs attention. A step that fails
|
|
*every* time is a loud, repeating log entry, which is the correct signal for "this recurring job is
|
|
broken" — distinct from "this machine did not converge."
|
|
|
|
**Runs do not stack.** If a run is still going when the next is due, the host skips the due run rather
|
|
than starting a second copy, and logs the skip. A slow nightly job that occasionally overruns must not
|
|
spawn a growing pile of concurrent containers competing for the same account and volumes — the failure
|
|
mode that made the old timers dangerous.
|
|
|
|
**Each run is independent and idempotent by the module's own code.** The mesh guarantees only *the
|
|
container is run on the cadence*; that a run does the right thing when the previous one half-finished
|
|
is the module's contract, the same discipline a run-once seed owes
|
|
([ADR 0052](0052-a-step-that-runs-once-before-a-container.md),
|
|
[04-ISSUES/035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md)).
|
|
|
|
This is a decision and not a patch because it settles **what a module may say about running its own
|
|
code on a cadence**, which every recurring provider job — a sync, a poll, a roll-up — now and later
|
|
depends on, and because it draws the line 0052 left: the recurring twin of run-once, with the gating
|
|
rule deliberately reversed so a routine job's failure is never a machine's failure.
|
|
|
|
### The security bound, stated plainly
|
|
|
|
**The host gains no new action and no new shape.** `schedule` is a string modifier on the `container`
|
|
shape that already exists. A scheduled container is strictly *less* powerful than an `action`: it runs
|
|
a digest-pinned image under an account the mesh scoped, which is exactly what `container` already
|
|
grants from the link, and it cannot run an arbitrary host command. A compromised control plane can
|
|
express nothing through `schedule` it could not already express by declaring a `container` — the cron
|
|
string only says *how often*, not *what*. The dangerous expansion of option 1 — a command the host
|
|
runs on the machine on a timer — is not made.
|
|
|
|
### How each claim is checked
|
|
|
|
- **A scheduled step runs when the schedule is due.** A host unit test installs a container with a
|
|
schedule that is due immediately (or advances a injected clock to when it is due) and asserts the
|
|
host ran it to completion; a sibling test with a schedule not yet due asserts it has not run.
|
|
- **Installing it does not run it, and the node is current without a run.** A host unit test applies a
|
|
scheduled container and asserts the apply reports current *before* any run has fired — the schedule
|
|
is state that is present, not a step that gated.
|
|
- **A failed run does not fail the apply or the node.** A host unit test fires a scheduled container
|
|
that exits non-zero and asserts the failure is recorded against the module, the apply is not failed,
|
|
and the node stays current — the mirror of the run-once test where a non-zero exit *does* halt.
|
|
- **Runs do not stack.** A host unit test fires a scheduled container whose run outlasts its next due
|
|
time and asserts the host skipped the due run and logged the skip, rather than starting a second
|
|
container.
|
|
- **The vocabulary carries the field end to end.** A control-plane unit test resolves a module whose
|
|
manifest sets `schedule` on a container and asserts the rendered host declaration carries the field;
|
|
the manifest parser refuses a `schedule` that is not a valid cron expression, and refuses a container
|
|
that is both `run-once` and `schedule` (a step is one or the other, never both).
|
|
- **A real module recurs in the lab.** A converted module declaring a scheduled step (kometa's library
|
|
sync, or a poller) is assigned in a lab scenario, and the scenario asserts the scheduled container
|
|
fires on its cadence and its account and volumes are the module's — the end-to-end proof this record
|
|
owes, as 0052 owed its run-once lab proof.
|
|
|
|
## Consequences
|
|
|
|
- **The recurring providers gain a home for their cadence.** kometa reconciles on its schedule; a
|
|
poller polls; a roll-up rolls up — each a scheduled step in its own runtime image, under its own
|
|
account, on the cron it declares.
|
|
- **The general lifecycle hook is still deferred.** Only *run once before* (0052) and *run on a
|
|
cadence* (this) are bought. `post-start`, `pre-remove` and the build/publish phases remain unbuilt,
|
|
decided when a real need arrives, not speculatively — the restraint that kept both from being the old
|
|
engine.
|
|
- **A recurring job's failure is loud but not fatal.** The node stays current while a scheduled step
|
|
fails and retries; a step that fails forever is a repeating log entry, not a machine marked broken.
|
|
This is the correct separation — a machine's convergence and a job's success are different questions —
|
|
and it is why this could not simply be `run-once` without the schedule.
|
|
- **The host vocabulary did not grow, again, and that is the point.** The cost of a scheduled step is
|
|
one string field and a cron loop in the container applier, not a new shape and not a new action. The
|
|
mesh expresses cadence; the module runs its own code, where it already runs it.
|
|
- **`run-once` and `schedule` are exclusive and complete for now.** A container runs once and gates, or
|
|
runs on a cadence and does not, or runs and stays up (a service). A manifest that asks for two of
|
|
these at once is refused, because the three are distinct answers to "how does this container run."
|
|
|
|
## References
|
|
|
|
- [ADR 0052](0052-a-step-that-runs-once-before-a-container.md) — the run-once step; this is its
|
|
recurring twin, reusing the `container`-modifier shape and inheriting its security bound, and
|
|
reversing its gating rule
|
|
- [04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md) — a
|
|
module cannot run its own code at a lifecycle phase; 0052 closed the once-before-start face, this
|
|
closes the recurring face
|
|
- [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) — a module runs
|
|
its code as its own process under its own account; a scheduled step is that process, run on a cadence
|
|
- [ADR 0005](0005-the-node-host.md) — the host's finite vocabulary, and why a host command (even on a
|
|
timer) is refused from the link
|
|
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — a container is pinned by digest; a
|
|
scheduled container no differently
|
|
- [ADR 0018](0018-a-picture-is-read-from-what-runs.md) — what was applied is recorded after it works;
|
|
the installed schedule is state, a fired run is an event
|
|
- mesh-control `feat/schedule-container`, mesh-host `feat/apply-schedule`, and the converted module
|
|
that first declares a scheduled step — the vocabulary, the apply support, and the recurring proof
|