openmp
Differences
This shows you the differences between two versions of the page.
| Both sides previous revisionPrevious revisionNext revision | Previous revision | ||
| openmp [June 11, 2026 at 11:03] – Ivan Janevski | openmp [August 22, 2026 at 15:22] (current) – external edit 127.0.0.1 | ||
|---|---|---|---|
| Line 1: | Line 1: | ||
| # OpenMP | # OpenMP | ||
| - | **OpenMP** is a shared-memory parallelism API for C, C++, and Fortran. | ||
| - | If you have a loop that takes too long and you want it to use all the cores on your machine instead of just one, OpenMP is usually the shortest path there. It works via compiler directives (`#pragma omp` in C/C++), a small runtime library (`libomp`), and a set of environment variables. In contrast to distributed-memory models like [[mpi|MPI]], where each process has its own memory, OpenMP threads all share the same address space. You get parallelism without touching your data layout. | + | **OpenMP** is a shared-memory parallelism API for C, C++, and Fortran. Add a `#pragma omp` directive before |
| - | The execution model is **fork-join**: | + | The execution model is **fork-join**: |
| - | + | ||
| - | ```c | + | |
| - | #pragma omp parallel | + | |
| - | { | + | |
| - | int tid = omp_get_thread_num(); | + | |
| - | int nthreads = omp_get_num_threads(); | + | |
| - | printf(" | + | |
| - | } | + | |
| - | ``` | + | |
| - | + | ||
| - | The thread | + | |
| - | + | ||
| - | Adding more threads does not always mean proportionally faster code. [[amdahls-law|Amdahl' | + | |
| - | + | ||
| - | ## Practice | + | |
| - | + | ||
| - | Compile the hello-world above and run it: | + | |
| - | + | ||
| - | ```bash | + | |
| - | $ gcc -fopenmp -o hello hello.c | + | |
| - | $ OMP_NUM_THREADS=4 ./hello | + | |
| - | thread 2 of 4 | + | |
| - | thread 0 of 4 | + | |
| - | thread 3 of 4 | + | |
| - | thread 1 of 4 | + | |
| - | ``` | + | |
| - | + | ||
| - | The output order is non-deterministic — run it a few times and you will get different permutations. Now try dropping `-fopenmp`: | + | |
| - | + | ||
| - | ```bash | + | |
| - | $ gcc -o hello hello.c | + | |
| - | $ ./hello | + | |
| - | thread 0 of 1 | + | |
| - | ``` | + | |
| - | + | ||
| - | Without the flag, every `#pragma omp` directive is silently ignored and the program runs as a single thread. This is the most useful OpenMP debugging technique: if your parallel output looks wrong, remove `-fopenmp` and check whether the serial output is correct first. | + | |
| - | + | ||
| - | Here is a more realistic example: parallelising a sum of a billion integers with a single pragma. | + | |
| ```c | ```c | ||
| // compile: gcc -O2 -fopenmp -o sum sum.c | // compile: gcc -O2 -fopenmp -o sum sum.c | ||
| // run: OMP_NUM_THREADS=4 ./sum | // run: OMP_NUM_THREADS=4 ./sum | ||
| - | // description: | + | // description: |
| #include < | #include < | ||
| Line 63: | Line 24: | ||
| ``` | ``` | ||
| - | On a 4-core machine this runs roughly 4× faster with `-fopenmp` than without. | + | On a 4-core machine this runs roughly 4× faster with `-fopenmp` than without. |
| - | + | ||
| - | Try varying `OMP_NUM_THREADS` from 1 up to your core count and beyond. Speedup will plateau or even decline past the hardware thread count — that is thread management overhead and Amdahl' | + | |
| ## Concepts | ## Concepts | ||
| - | 1. [[data-sharing-openmp|Data sharing]] | + | 1. [[openmp-data-sharing|Data sharing]] |
| - | 2. [[parallel-loops-openmp|Parallel loops]] | + | 2. [[openmp-parallel-loops|Parallel loops]] |
| - | 3. [[collapse-openmp|Collapse]] | + | 3. [[openmp-collapse|Collapse]] |
| - | 4. [[reduction-openmp|Reduction]] | + | 4. [[openmp-reduction|Reduction]] |
| - | 5. [[scheduling-openmp|Scheduling]] | + | 5. [[openmp-scheduling|Scheduling]] |
| - | 6. [[simd-openmp|SIMD]] | + | 6. [[openmp-simd|SIMD]] |
| - | 7. [[tasks-openmp|Tasks]] | + | 7. [[openmp-tasks|Tasks]] |
| - | 8. [[single-openmp|Single]] | + | 8. [[openmp-single|Single]] |
| - | 9. [[master-openmp|Master]] | + | 9. [[openmp-master|Master]] |
| - | 10. [[sections-openmp|Sections]] | + | 10. [[openmp-sections|Sections]] |
| - | 11. [[barrier-openmp|Barrier]] | + | 11. [[openmp-barrier|Barrier]] |
| - | 12. [[nowait-openmp|Nowait]] | + | 12. [[openmp-nowait|Nowait]] |
| - | 13. [[critical-sections-openmp|Critical sections]] | + | 13. [[openmp-critical-sections|Critical sections]] |
| - | 14. [[atomic-openmp|Atomic]] | + | 14. [[openmp-atomic|Atomic]] |
| - | 15. [[flush-openmp|Flush]] | + | 15. [[openmp-flush|Flush]] |
| - | 16. [[false-sharing-openmp|False sharing]] | + | 16. [[openmp-false-sharing|False sharing]] |
| - | 17. [[thread-affinity-openmp|Thread affinity]] | + | 17. [[openmp-thread-affinity|Thread affinity]] |
| - | 18. [[time-measurement-openmp|Time measurement]] | + | 18. [[openmp-time-measurement|Time measurement]] |
| - | + | 19. [[openmp-overview|Overview]] | |
| - | ## Overview | + | |
| - | + | ||
| - | ### Directives | + | |
| - | + | ||
| - | ```c | + | |
| - | #pragma omp parallel | + | |
| - | #pragma omp parallel for // distribute loop iterations across the team | + | |
| - | #pragma omp parallel for reduction(+: | + | |
| - | #pragma omp parallel sections | + | |
| - | #pragma omp section | + | |
| - | #pragma omp single | + | |
| - | #pragma omp master | + | |
| - | #pragma omp task // package work for any idle thread to execute | + | |
| - | #pragma omp taskwait | + | |
| - | #pragma omp barrier | + | |
| - | #pragma omp critical | + | |
| - | #pragma omp atomic | + | |
| - | #pragma omp simd // assert the loop is safe to vectorise | + | |
| - | #pragma omp flush // enforce memory visibility across threads | + | |
| - | ``` | + | |
| - | + | ||
| - | ### Functions | + | |
| - | + | ||
| - | ```c | + | |
| - | omp_get_thread_num() | + | |
| - | omp_get_num_threads() | + | |
| - | omp_get_max_threads() | + | |
| - | omp_set_num_threads(n) | + | |
| - | omp_get_num_procs() | + | |
| - | omp_get_wtime() | + | |
| - | omp_in_parallel() | + | |
| - | ``` | + | |
| - | + | ||
| - | ### Environment variables | + | |
| - | + | ||
| - | ^ Variable ^ Default ^ Description ^ | + | |
| - | | `OMP_NUM_THREADS` | core count | Number of threads to use in each parallel region | | + | |
| - | | `OMP_SCHEDULE` | `static` | Default schedule kind and optional chunk size, e.g. `dynamic,4` | | + | |
| - | | `OMP_PROC_BIND` | `false` | Thread-to-core affinity policy: `close`, `spread`, or `master` | | + | |
| - | | `OMP_PLACES` | (unset) | Placement units for affinity: `cores`, `threads`, or `sockets` | | + | |
| - | | `OMP_MAX_ACTIVE_LEVELS` | `1` | Maximum nesting depth of simultaneously active parallel regions | | + | |
| - | | `OMP_DISPLAY_ENV` | `false` | Print OpenMP version and active settings at startup: `TRUE` or `VERBOSE` | + | |
openmp.1781175830.md.gz · Last modified: by Ivan Janevski
