openmp
Differences
This shows you the differences between two versions of the page.
| Both sides previous revisionPrevious revisionNext revision | Previous revision | ||
| openmp [June 11, 2026 at 10:22] – external edit 127.0.0.1 | openmp [August 22, 2026 at 15:22] (current) – external edit 127.0.0.1 | ||
|---|---|---|---|
| Line 1: | Line 1: | ||
| # OpenMP | # OpenMP | ||
| - | **OpenMP** is a shared-memory parallelism API for C, C++, and Fortran. If you have a loop that takes too long and you want it to use all the cores on your machine instead of just one, OpenMP is usually the shortest path there. It works via compiler directives (`#pragma omp` in C/C++), a small runtime library (`libomp`), and a set of environment variables. In contrast to distributed-memory models like [[mpi|MPI]], | ||
| - | The execution model is **fork-join**: the program starts as a single thread. When it hits a `#pragma omp parallel` block, it forks into a team of worker threads that all execute the block concurrently, | + | **OpenMP** is a shared-memory parallelism API for C, C++, and Fortran. Add a `#pragma omp` directive before a loop or block to parallelize |
| - | ```c | + | The execution model is **fork-join**: |
| - | #pragma omp parallel | + | |
| - | { | + | |
| - | int tid = omp_get_thread_num(); | + | |
| - | int nthreads = omp_get_num_threads(); | + | |
| - | printf(" | + | |
| - | } | + | |
| - | ``` | + | |
| - | + | ||
| - | The thread | + | |
| - | + | ||
| - | Adding more threads does not always mean proportionally faster code. [[amdahls-law|Amdahl' | + | |
| - | + | ||
| - | ## Practice | + | |
| - | + | ||
| - | Compile the hello-world above and run it: | + | |
| - | + | ||
| - | ```bash | + | |
| - | $ gcc -fopenmp -o hello hello.c | + | |
| - | $ OMP_NUM_THREADS=4 ./hello | + | |
| - | thread 2 of 4 | + | |
| - | thread 0 of 4 | + | |
| - | thread 3 of 4 | + | |
| - | thread 1 of 4 | + | |
| - | ``` | + | |
| - | + | ||
| - | The output order is non-deterministic — run it a few times and you will get different permutations. Now try dropping `-fopenmp`: | + | |
| - | + | ||
| - | ```bash | + | |
| - | $ gcc -o hello hello.c | + | |
| - | $ ./hello | + | |
| - | thread 0 of 1 | + | |
| - | ``` | + | |
| - | + | ||
| - | Without the flag, every `#pragma omp` directive is silently ignored and the program runs as a single thread. This is the most useful OpenMP debugging technique: if your parallel output looks wrong, remove `-fopenmp` and check whether the serial output is correct first. | + | |
| - | + | ||
| - | Here is a more realistic example: parallelising a sum of a billion integers with a single pragma. | + | |
| ```c | ```c | ||
| // compile: gcc -O2 -fopenmp -o sum sum.c | // compile: gcc -O2 -fopenmp -o sum sum.c | ||
| // run: OMP_NUM_THREADS=4 ./sum | // run: OMP_NUM_THREADS=4 ./sum | ||
| - | // description: | + | // description: |
| #include < | #include < | ||
| Line 61: | Line 24: | ||
| ``` | ``` | ||
| - | On a 4-core machine this runs roughly 4× faster with `-fopenmp` than without. | + | On a 4-core machine this runs roughly 4× faster with `-fopenmp` than without. |
| - | + | ||
| - | Try varying `OMP_NUM_THREADS` from 1 up to your core count and beyond. Speedup will plateau or even decline past the hardware thread count — that is thread management overhead and Amdahl' | + | |
| ## Concepts | ## Concepts | ||
| - | 1. [[data-sharing-openmp|Data sharing]] | + | 1. [[openmp-data-sharing|Data sharing]] |
| - | 2. [[parallel-loops-openmp|Parallel loops]] | + | 2. [[openmp-parallel-loops|Parallel loops]] |
| - | 3. [[collapse-openmp|Collapse]] | + | 3. [[openmp-collapse|Collapse]] |
| - | 4. [[reduction-openmp|Reduction]] | + | 4. [[openmp-reduction|Reduction]] |
| - | 5. [[scheduling-openmp|Scheduling]] | + | 5. [[openmp-scheduling|Scheduling]] |
| - | 6. [[simd-openmp|SIMD]] | + | 6. [[openmp-simd|SIMD]] |
| - | 7. [[tasks-openmp|Tasks]] | + | 7. [[openmp-tasks|Tasks]] |
| - | 8. [[single-openmp|Single]] | + | 8. [[openmp-single|Single]] |
| - | 9. [[master-openmp|Master]] | + | 9. [[openmp-master|Master]] |
| - | 10. [[sections-openmp|Sections]] | + | 10. [[openmp-sections|Sections]] |
| - | 11. [[barrier-openmp|Barrier]] | + | 11. [[openmp-barrier|Barrier]] |
| - | 12. [[nowait-openmp|Nowait]] | + | 12. [[openmp-nowait|Nowait]] |
| - | 13. [[critical-sections-openmp|Critical sections]] | + | 13. [[openmp-critical-sections|Critical sections]] |
| - | 14. [[atomic-openmp|Atomic]] | + | 14. [[openmp-atomic|Atomic]] |
| - | 15. [[flush-openmp|Flush]] | + | 15. [[openmp-flush|Flush]] |
| - | 16. [[false-sharing-openmp|False sharing]] | + | 16. [[openmp-false-sharing|False sharing]] |
| - | 17. [[thread-affinity-openmp|Thread affinity]] | + | 17. [[openmp-thread-affinity|Thread affinity]] |
| - | 18. [[time-measurement-openmp|Time measurement]] | + | 18. [[openmp-time-measurement|Time measurement]] |
| - | + | 19. [[openmp-overview|Overview]] | |
| - | ## Overview | + | |
| - | + | ||
| - | ### Directives | + | |
| - | + | ||
| - | ```c | + | |
| - | #pragma omp parallel | + | |
| - | #pragma omp parallel for // distribute loop iterations across the team | + | |
| - | #pragma omp parallel for reduction(+: | + | |
| - | #pragma omp parallel sections | + | |
| - | #pragma omp section | + | |
| - | #pragma omp single | + | |
| - | #pragma omp master | + | |
| - | #pragma omp task // package work for any idle thread to execute | + | |
| - | #pragma omp taskwait | + | |
| - | #pragma omp barrier | + | |
| - | #pragma omp critical | + | |
| - | #pragma omp atomic | + | |
| - | #pragma omp simd // assert the loop is safe to vectorise | + | |
| - | #pragma omp flush // enforce memory visibility across threads | + | |
| - | ``` | + | |
| - | + | ||
| - | ### Functions | + | |
| - | + | ||
| - | ```c | + | |
| - | omp_get_thread_num() | + | |
| - | omp_get_num_threads() | + | |
| - | omp_get_max_threads() | + | |
| - | omp_set_num_threads(n) | + | |
| - | omp_get_num_procs() | + | |
| - | omp_get_wtime() | + | |
| - | omp_in_parallel() | + | |
| - | ``` | + | |
| - | + | ||
| - | ### Environment variables | + | |
| - | + | ||
| - | ^ Variable ^ Default ^ Description ^ | + | |
| - | | `OMP_NUM_THREADS` | core count | Number of threads to use in each parallel region | | + | |
| - | | `OMP_SCHEDULE` | `static` | Default schedule kind and optional chunk size, e.g. `dynamic,4` | | + | |
| - | | `OMP_PROC_BIND` | `false` | Thread-to-core affinity policy: `close`, `spread`, or `master` | | + | |
| - | | `OMP_PLACES` | (unset) | Placement units for affinity: `cores`, `threads`, or `sockets` | | + | |
| - | | `OMP_MAX_ACTIVE_LEVELS` | `1` | Maximum nesting depth of simultaneously active parallel regions | | + | |
| - | | `OMP_DISPLAY_ENV` | `false` | Print OpenMP version and active settings at startup: `TRUE` or `VERBOSE` | + | |
openmp.1781173372.md.gz · Last modified: by 127.0.0.1
