Run with OpenMP¶
Build the adc_cpp C++ facade against a Kokkos OpenMP execution space and run the
compiled tests multi-threaded on CPU. Use this when the Kokkos Serial backend is too slow on a
single thread and you want to use the cores of one machine.
The backend is fixed at compile time, not chosen at runtime, so changing it means reconfiguring
and recompiling. This recipe assumes you already build adc_cpp with the Kokkos Serial backend;
see Installation if you do not. For the full list of
parallel configurations, see the parallel backends page.
A multi-thread (Kokkos OpenMP) Python module is available too: build it with the
python-parallel CMake preset, then set the thread count from Python with pops.set_threads(n)
called right after import pops. The remaining caveats are real: there is no nvcc/CUDA Python
module (GPU runs are C++-only), and the Python pops module does not exercise the MPI code paths
(MPI is validated through the C++/ctest path).
Steps¶
Install a Kokkos with
Kokkos_ENABLE_OPENMP=ON. Point a path variable at that install. ReplaceKOKKOS_OPENMP_PREFIXwith the absolute path to the Kokkos OpenMP install prefix.export KOKKOS_OPENMP_PREFIX=/path/to/kokkos-openmp
Configure the build against that Kokkos. The
-DPOPS_USE_KOKKOS=ONflag is the same as for the Serial build; only the Kokkos install pointed at by-DKokkos_ROOTchanges.cmake -S . -B build-kokkos-omp -DCMAKE_BUILD_TYPE=Release -DPOPS_USE_KOKKOS=ON -DKokkos_ROOT="$KOKKOS_OPENMP_PREFIX"
Compile the facade and the tests.
cmake --build build-kokkos-omp -j
Run the tests, bounding the thread count with
OMP_NUM_THREADS. Set it to the number of cores you want to use.OMP_NUM_THREADS=4 ctest --test-dir build-kokkos-omp --output-on-failure
With the OpenMP execution space, for_each_cell becomes a multi-thread parallel_for. The sum
reduction reassociates floating-point addition per tile, so a multi-thread total is deterministic
for a fixed thread count but not bit-identical to a lexicographic sum; the max reduction stays exact.
Threading from Python (set_threads)¶
The Python module sets its thread count with pops.set_threads(n), which writes OMP_NUM_THREADS
and KOKKOS_NUM_THREADS for you instead of exporting them in the shell. It is pure Python: no
C++ call, so the value must be in place before Kokkos initializes.
import pops
pops.set_threads(8) # or pops.set_threads() for all cores (os.cpu_count())
print(pops.parallel_info()) # {'has_kokkos': True, 'omp_num_threads': '8', 'first_system_built': False}
Three rules:
Use a module built against a Kokkos OpenMP execution space (
python-parallelpreset). The conda-forge Kokkos is often Serial-only; see Kokkos OpenMP and the Installation Threads section.Call
set_threadsright afterimport popsand before the first bind/runtime object. Kokkos reads the thread count once at that first runtime allocation, so a later call cannot change it.A Serial-only module or a late call only emits a
RuntimeWarningand is ignored; it never raises. Confirm the state withpops.has_kokkos(),pops.parallel_info(), orpops.doctor().
The floating-point note above applies to the Python path too: the per-tile sum reduction is
deterministic for a fixed thread count but not bit-identical to the serial sum; the max is exact.
POPS_FOREACH_SERIAL_THRESHOLD keeps small grids on a serial loop, so small problems may not speed
up regardless of the thread count.
Next steps¶
To add distributed ranks on top of OpenMP, combine
-DPOPS_USE_MPI=ONwith the same OpenMP Kokkos install, described on the parallel backends page.To check which execution space a build actually runs, read Checking your backend.