Program scheduler¶
Spec 3 unifies cadence into a single Program scheduler: any Program operation can carry a schedule, replacing the scattered block stride / substeps / source frequency knobs.
Status
The schedule AUTHORING surface (the vocabulary below, policy chaining, recording a schedule on a
Program node, the cacheable-capability validation) and the CODEGEN that lowers every kind/policy to
C++ are available in pops.time (ADC-458). A scheduled node lowers to a due-test guard plus its
policy branch around the statements it emits (the cache lives in the C++ CacheManager). Two cases
still refuse to lower, loudly: on_end() (a compiled sim.step(dt) loop carries no end-of-run
signal, so the .so cannot know the last step – use an on_end host hook) and a when(cond) over a
bare Python callable (only a Program Bool predicate lowers). The cache cadence RUNTIME in a stepping
.so is exercised on ROMEO; the CacheManager and its checkpoint round-trip are unit-tested by
tests/test_cache_manager.cpp and tests/test_checkpoint_cache.cpp. The cache state is checkpointed
(below).
Schedules¶
A schedule decides when a node is due:
Schedule |
Meaning |
|---|---|
|
due every step (the default) |
|
due every N macro-steps |
|
due when a runtime condition holds |
|
due at the first / last step |
|
structured sub-cycling of a block |
and a policy decides what happens when it is NOT due:
Policy |
Behaviour |
|---|---|
|
always recompute (default) |
|
reuse the last cached value |
|
produce nothing (diagnostics only) |
|
return zero (optional contributions) |
|
accumulate the skipped dt and apply with |
|
error if not due |
With a variable step_cfl, accumulate_dt must use the real sum of the skipped dt, not
N * dt_current.
Cacheable capabilities¶
hold is only valid on a cacheable operator. An operator declares this:
m.operator_capabilities("fields_from_state", cacheable=True, stale_allowed=True)
m.operator_capabilities("explicit_rate", cacheable=False, requires_fresh_inputs=True)
so requesting every(10).hold() on a non-cacheable operator is an error:
operator 'explicit_rate' is not cacheable; cannot use schedule hold.
Lowering¶
A scheduled node lowers around the C++ its operation emits. The due test depends on the kind:
every(N) -> ctx.cache_should_update(id, N); on_start() -> ctx.macro_step() == 0; when(cond)
-> the lowered Program Bool predicate; subcycle(count, dt) -> a for sub-loop over the sub-dt. The
policy decides the not-due branch: recompute / skip run the body only when due (skip keeps the
stale value, the cacheable contract); zero sets the output to zero off-cadence; hold stores the
output when due and restores it otherwise; accumulate_dt accumulates the actual skipped dt and reads
eff_dt = sum(dt_skipped) on the due step (never N * dt_current); error fails loud if read
off-cadence. A field solve caches the System aux; any other node caches its own named scratch.
Cache and checkpoint¶
The CacheManager (include/pops/runtime/program/cache_manager.hpp) stores per node the cached value
(aux or named scratch), the last update step, the accumulated dt and a validity flag. It is typed,
C++-allocated and Kokkos/MPI-safe. The cache is owned by the System (System::program_cache()), not
the .so step closure, so the EXISTING checkpoint reaches it: sim.checkpoint gathers each valid slot
exactly the way it gathers the multistep history rings (program_cache_global mirrors
history_global), and sim.restart scatters them back (restore_program_cache). A
(run, checkpoint, restart, continue) run is then bit-for-bit identical to a continuous run – the held
node resumes on its cadence instead of cold-starting. Two guards fail loud at restart: a restart against
a DIFFERENT compiled Program is rejected by the program-hash guard
(checkpoint was created with a different compiled Program hash), and a checkpoint that lists a cached
node but lost its value array names the node
(checkpoint missing cached value for scheduled node '<name>'). A checkpoint with no cached node
restores as before (back-compatible). The cache value MultiFab is serialized through the same
gather_global / write_state machinery as the block state, so the round-trip is MPI-safe and
bit-identical under np>1; the full compiled-.so held-schedule continuous == restart run is
Kokkos-only AOT (ROMEO).
Per-block step policy¶
The older per-block step policy still runs alongside the scheduler: a block advances substeps
sub-steps of dt/substeps, or 1 macro-step out of stride. See Time programs.