The recurring twin of run-once (0052): a schedule modifier on the container shape, reusing its security bound (no new host action/shape, strictly less than an action) and reversing its gating rule — a scheduled step runs after convergence, does not gate the apply, and a failed run is logged, not fatal. Unblocks kometa's sync and pollers. Accepted per direction to build the primitive now. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
13 KiB
topic, status, date, deciders, reconstructed, extends
| topic | status | date | deciders | reconstructed | extends |
|---|---|---|---|---|---|
| what runs on it | accepted | 2026-09-06 | jochen | false | 0052-a-step-that-runs-once-before-a-container.md |
53. A scheduled step is a container run on a recurring schedule
Context
ADR 0052 gave the mesh a step that runs once; a real class of modules needs one that runs again and again. kometa reconciles a media library against its lists on a timer; a ticketing integration polls its source for new work every few minutes; a backup, a cache warm, a metrics roll-up all recur. The host's vocabulary describes state — a file, a directory, a container that should be running — and 0052 added a step that happens once and is done. Neither says this should happen every night at 3, forever (04-ISSUES/037 named the lifecycle gap; 0052 closed the once-before-start face of it and left the recurring face open).
The modules that need it already run their own code. As with run-once, the ground has shifted since the old mesh's flaky hooks: a module with tools or events runs a process of its own — a container carrying the module's compiled code under the single scoped account the mesh gave it (ADR 0047). The work that must recur is already in that image, under that account. The mesh does not need a way to run a module's code on a timer — it has the code and the account. It needs a way to say run this container again on this cadence.
The shape of the answer is already decided, one modifier over. 0052 rejected a host run command,
per-phase hooks, and a distinct one-shot resource type, and adopted a modifier on the container
shape the host already has, because a one-shot is a container that happens to exit. A scheduled step
is the same container that happens to exit — run repeatedly. Reusing that shape keeps the host
vocabulary flat and inherits 0052's security bound whole. The only thing 0052's run-once does not
carry is when to run it again.
One difference from run-once changes a rule, and it is the reason this is its own record. A run-once step gates the apply: it is placed before the container that depends on it, and a failure halts everything after it, because "seed the store before the broker starts" is a correctness precondition (ADR 0052). A scheduled step is the opposite: it runs after the machine is up and converged, on its own clock, and a single failed run is an ordinary operational event — the next run comes anyway. A scheduled step that halted the apply, or that a failed run marked the node not-current over, would make a routine poll into a reason the whole machine reads as broken. So the gating rule 0052 established is exactly the rule this record must not inherit.
Considered Options
-
A host
cron/timershape — the host installs a system timer that runs a command. The direct reading. Rejected, for the reason 0052 rejected a hostrunshape: it widens what a compromised control plane can express toward "run this command on the machine, forever," which is the largest widening there is, and it isaction-from-the-link by another name (ADR 0005). A recurring command is worse than a one-off, because it persists. -
Per-phase lifecycle hooks —
on-schedulejoiningpre-start/post-startas named code the mesh runs. Rejected for now, as in 0052: it is the flaky hook engine the issue warns against, and the three modules that need this need one thing — run this container on a cadence — which a narrow modifier expresses without deciding a whole hook vocabulary. -
A distinct
scheduledresource type, sibling tocontainer. Rejected, as 0052 rejected a distinct one-shot type: it spends a whole new host shape (aType, a struct, an applier, a place in every host's shape list) on something acontaineralready almost is — a scheduled task is a container (pinned image, account, volumes, environment) that runs on a clock. -
A modifier on the existing
containershape:schedule, a cron expression the host runs the container on. Adopted. It reuses the shape the host has, adds no new host action, and sits besiderun-onceas its recurring twin — the same container, exited, run again.
Decision
A scheduled step is an ordinary container, marked with a schedule. The manifest sets
schedule: "<cron>" on a container resource — a standard five-field cron expression. Everything else
about it is a container as before: a digest-pinned image
(ADR 0006), volumes, an environment, and for a module's
own code the same scoped account its runtime already holds
(ADR 0047). The mesh adds no
way to run code; it marks a container the host must run on that cadence, each time to completion,
rather than start once and leave running (a service) or run once and gate (a run-once step).
A run is fired by the clock, not by the apply, and does not gate it. Applying the declaration installs the schedule; it does not run the step. The machine converges — reports applied and current — as soon as the schedule is installed, exactly as it does for a service that is running. Thereafter the host fires the container when the cron expression is due. This is the deliberate inversion of 0052: a scheduled step is downstream of convergence, not a precondition of it.
A failed run is recorded and the next run still comes; it never marks the node not-current. A run that exits non-zero is logged against the module — visible, auditable — but it does not fail the apply, does not halt other resources, and does not flip the node's reported state. A poll that fails at 03:00 and succeeds at 03:05 is the system working, not a machine that needs attention. A step that fails every time is a loud, repeating log entry, which is the correct signal for "this recurring job is broken" — distinct from "this machine did not converge."
Runs do not stack. If a run is still going when the next is due, the host skips the due run rather than starting a second copy, and logs the skip. A slow nightly job that occasionally overruns must not spawn a growing pile of concurrent containers competing for the same account and volumes — the failure mode that made the old timers dangerous.
Each run is independent and idempotent by the module's own code. The mesh guarantees only the container is run on the cadence; that a run does the right thing when the previous one half-finished is the module's contract, the same discipline a run-once seed owes (ADR 0052, 04-ISSUES/035).
This is a decision and not a patch because it settles what a module may say about running its own code on a cadence, which every recurring provider job — a sync, a poll, a roll-up — now and later depends on, and because it draws the line 0052 left: the recurring twin of run-once, with the gating rule deliberately reversed so a routine job's failure is never a machine's failure.
The security bound, stated plainly
The host gains no new action and no new shape. schedule is a string modifier on the container
shape that already exists. A scheduled container is strictly less powerful than an action: it runs
a digest-pinned image under an account the mesh scoped, which is exactly what container already
grants from the link, and it cannot run an arbitrary host command. A compromised control plane can
express nothing through schedule it could not already express by declaring a container — the cron
string only says how often, not what. The dangerous expansion of option 1 — a command the host
runs on the machine on a timer — is not made.
How each claim is checked
- A scheduled step runs when the schedule is due. A host unit test installs a container with a schedule that is due immediately (or advances a injected clock to when it is due) and asserts the host ran it to completion; a sibling test with a schedule not yet due asserts it has not run.
- Installing it does not run it, and the node is current without a run. A host unit test applies a scheduled container and asserts the apply reports current before any run has fired — the schedule is state that is present, not a step that gated.
- A failed run does not fail the apply or the node. A host unit test fires a scheduled container that exits non-zero and asserts the failure is recorded against the module, the apply is not failed, and the node stays current — the mirror of the run-once test where a non-zero exit does halt.
- Runs do not stack. A host unit test fires a scheduled container whose run outlasts its next due time and asserts the host skipped the due run and logged the skip, rather than starting a second container.
- The vocabulary carries the field end to end. A control-plane unit test resolves a module whose
manifest sets
scheduleon a container and asserts the rendered host declaration carries the field; the manifest parser refuses aschedulethat is not a valid cron expression, and refuses a container that is bothrun-onceandschedule(a step is one or the other, never both). - A real module recurs in the lab. A converted module declaring a scheduled step (kometa's library sync, or a poller) is assigned in a lab scenario, and the scenario asserts the scheduled container fires on its cadence and its account and volumes are the module's — the end-to-end proof this record owes, as 0052 owed its run-once lab proof.
Consequences
- The recurring providers gain a home for their cadence. kometa reconciles on its schedule; a poller polls; a roll-up rolls up — each a scheduled step in its own runtime image, under its own account, on the cron it declares.
- The general lifecycle hook is still deferred. Only run once before (0052) and run on a
cadence (this) are bought.
post-start,pre-removeand the build/publish phases remain unbuilt, decided when a real need arrives, not speculatively — the restraint that kept both from being the old engine. - A recurring job's failure is loud but not fatal. The node stays current while a scheduled step
fails and retries; a step that fails forever is a repeating log entry, not a machine marked broken.
This is the correct separation — a machine's convergence and a job's success are different questions —
and it is why this could not simply be
run-oncewithout the schedule. - The host vocabulary did not grow, again, and that is the point. The cost of a scheduled step is one string field and a cron loop in the container applier, not a new shape and not a new action. The mesh expresses cadence; the module runs its own code, where it already runs it.
run-onceandscheduleare exclusive and complete for now. A container runs once and gates, or runs on a cadence and does not, or runs and stays up (a service). A manifest that asks for two of these at once is refused, because the three are distinct answers to "how does this container run."
References
- ADR 0052 — the run-once step; this is its
recurring twin, reusing the
container-modifier shape and inheriting its security bound, and reversing its gating rule - 04-ISSUES/037 — a module cannot run its own code at a lifecycle phase; 0052 closed the once-before-start face, this closes the recurring face
- ADR 0047 — a module runs its code as its own process under its own account; a scheduled step is that process, run on a cadence
- ADR 0005 — the host's finite vocabulary, and why a host command (even on a timer) is refused from the link
- ADR 0006 — a container is pinned by digest; a scheduled container no differently
- ADR 0018 — what was applied is recorded after it works; the installed schedule is state, a fired run is an event
- mesh-control
feat/schedule-container, mesh-hostfeat/apply-schedule, and the converted module that first declares a scheduled step — the vocabulary, the apply support, and the recurring proof