Site Tools


openmp

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
openmp [June 11, 2026 at 11:03] Ivan Janevskiopenmp [August 22, 2026 at 15:22] (current) – external edit 127.0.0.1
Line 1: Line 1:
 # OpenMP # OpenMP
-**OpenMP** is a shared-memory parallelism API for C, C++, and Fortran. 
  
-If you have a loop that takes too long and you want it to use all the cores on your machine instead of just one, OpenMP is usually the shortest path there. It works via compiler directives (`#pragma omp` in C/C++), a small runtime library (`libomp`), and set of environment variablesIn contrast to distributed-memory models like [[mpi|MPI]]where each process has its own memoryOpenMP threads all share the same address space — you get parallelism without touching your data layout.+**OpenMP** is a shared-memory parallelism API for CC++, and Fortran. Add a `#pragma ompdirective before loop or block to parallelize it across available coresThe compiler handles thread creation and synchronization; if it doesn't support OpenMPdirectives are silently ignored and the program runs seriallymaking debugging easy.
  
-The execution model is **fork-join**: the program starts as a single thread. When it hits a `#pragma omp parallel` blockit forks into a team of worker threads that all execute the block concurrently, then join back into one thread at the closing braceYou do not write any thread creation, management, or teardown code — the compiler inserts all of that. If the compiler does not support OpenMP, it silently ignores all `#pragma omp` directives and the program runs serially, which is useful for debugging: drop `-fopenmp` from the compile command and you get a clean serial build without changing a line of source. +The execution model is **fork-join**: a single thread forks into worker team at `#pragma omp parallel`, they execute concurrently, then join back. Thread count defaults to the number of logical cores, configurable via `OMP_NUM_THREADS=N` at runtime or `omp_set_num_threads(n)` in code. Compile with `-fopenmp` and include `<omp.h>` for the `omp_*` runtime functions.
- +
-```c +
-#pragma omp parallel +
-+
-    int tid = omp_get_thread_num(); +
-    int nthreads = omp_get_num_threads(); +
-    printf("thread %d of %d\n", tid, nthreads); +
-+
-``` +
- +
-The thread count defaults to the number of logical cores. You can override it at runtime with `OMP_NUM_THREADS=N ./progwithout recompiling, or with `omp_set_num_threads(n)` inside the program, or with a `num_threads(N)` clause directly on the pragma. Output order from a parallel region is non-deterministic — threads are scheduled by the OS, not by rank. Compile with `-fopenmp` and include `<omp.h>` for the `omp_*` runtime functions+
- +
-Adding more threads does not always mean proportionally faster code. [[amdahls-law|Amdahl's law]] says that if a fraction $s$ of the program is inherently serial, the maximum speedup is $1/s$ no matter how many threads you add. A loop that accounts for 80% of runtime can at best give 5× speedup. This is why reducing the serial fraction matters more than just adding more cores, and why eliminating overhead like false sharing, load imbalance, and unnecessary barriers compounds those gains. +
- +
-## Practice +
- +
-Compile the hello-world above and run it: +
- +
-```bash +
-$ gcc -fopenmp -o hello hello.c +
-$ OMP_NUM_THREADS=4 ./hello +
-thread 2 of 4 +
-thread 0 of 4 +
-thread 3 of 4 +
-thread 1 of 4 +
-``` +
- +
-The output order is non-deterministic — run it a few times and you will get different permutations. Now try dropping `-fopenmp`: +
- +
-```bash +
-$ gcc -o hello hello.c +
-$ ./hello +
-thread 0 of 1 +
-``` +
- +
-Without the flag, every `#pragma omp` directive is silently ignored and the program runs as a single thread. This is the most useful OpenMP debugging technique: if your parallel output looks wrong, remove `-fopenmp` and check whether the serial output is correct first. +
- +
-Here is a more realistic example: parallelising a sum of a billion integers with a single pragma.+
  
 ```c ```c
 // compile: gcc -O2 -fopenmp -o sum sum.c // compile: gcc -O2 -fopenmp -o sum sum.c
 // run: OMP_NUM_THREADS=4 ./sum // run: OMP_NUM_THREADS=4 ./sum
-// description: parallel reduction; try with -fopenmp omitted for serial baseline+// description: parallel reduction using a single pragma
  
 #include <omp.h> #include <omp.h>
Line 63: Line 24:
 ``` ```
  
-On a 4-core machine this runs roughly 4× faster with `-fopenmp` than without. The only additions to the loop are the pragma line and the `reduction(+:sum)` clause. Without the clause you would get a data race and a wrong answer — the clause is what makes it correct. When it works this cleanly, it genuinely feels like cheating. +On a 4-core machine this runs roughly 4× faster with `-fopenmp` than without. [[amdahls-law|Amdahl's law]] limits speedup when part of the program is inherently serial; reducing overhead like false sharing and load imbalance matters more than just adding threads.
- +
-Try varying `OMP_NUM_THREADS` from 1 up to your core count and beyond. Speedup will plateau or even decline past the hardware thread count — that is thread management overhead and Amdahl's law at work.+
  
 ## Concepts ## Concepts
  
- 1. [[data-sharing-openmp|Data sharing]] + 1. [[openmp-data-sharing|Data sharing]] 
- 2. [[parallel-loops-openmp|Parallel loops]] + 2. [[openmp-parallel-loops|Parallel loops]] 
- 3. [[collapse-openmp|Collapse]] + 3. [[openmp-collapse|Collapse]] 
- 4. [[reduction-openmp|Reduction]] + 4. [[openmp-reduction|Reduction]] 
- 5. [[scheduling-openmp|Scheduling]] + 5. [[openmp-scheduling|Scheduling]] 
- 6. [[simd-openmp|SIMD]] + 6. [[openmp-simd|SIMD]] 
- 7. [[tasks-openmp|Tasks]] + 7. [[openmp-tasks|Tasks]] 
- 8. [[single-openmp|Single]] + 8. [[openmp-single|Single]] 
- 9. [[master-openmp|Master]] + 9. [[openmp-master|Master]] 
- 10. [[sections-openmp|Sections]] + 10. [[openmp-sections|Sections]] 
- 11. [[barrier-openmp|Barrier]] + 11. [[openmp-barrier|Barrier]] 
- 12. [[nowait-openmp|Nowait]] + 12. [[openmp-nowait|Nowait]] 
- 13. [[critical-sections-openmp|Critical sections]] + 13. [[openmp-critical-sections|Critical sections]] 
- 14. [[atomic-openmp|Atomic]] + 14. [[openmp-atomic|Atomic]] 
- 15. [[flush-openmp|Flush]] + 15. [[openmp-flush|Flush]] 
- 16. [[false-sharing-openmp|False sharing]] + 16. [[openmp-false-sharing|False sharing]] 
- 17. [[thread-affinity-openmp|Thread affinity]] + 17. [[openmp-thread-affinity|Thread affinity]] 
- 18. [[time-measurement-openmp|Time measurement]] + 18. [[openmp-time-measurement|Time measurement]] 
- + 19[[openmp-overview|Overview]]
-## Overview +
- +
-### Directives +
- +
-```c +
-#pragma omp parallel                         // fork a team of threads; join at closing brace +
-#pragma omp parallel for                     // distribute loop iterations across the team +
-#pragma omp parallel for reduction(+:s)     // loop with a parallel reduction +
-#pragma omp parallel sections                // distribute independent blocks across the team +
-#pragma omp section                          // one block inside a sections region +
-#pragma omp single                           // one thread runs the block; others wait at end +
-#pragma omp master                           // thread 0 only; no implicit barrier +
-#pragma omp task                             // package work for any idle thread to execute +
-#pragma omp taskwait                         // wait for all child tasks to finish +
-#pragma omp barrier                          // all threads wait until every thread arrives +
-#pragma omp critical                         // mutual exclusion — one thread at a time +
-#pragma omp atomic                           // single hardware-atomic read-modify-write +
-#pragma omp simd                             // assert the loop is safe to vectorise +
-#pragma omp flush                            // enforce memory visibility across threads +
-``` +
- +
-### Functions +
- +
-```c +
-omp_get_thread_num()      // ID of the calling thread (0 … N-1) +
-omp_get_num_threads()     // number of threads in the current team +
-omp_get_max_threads()     // threads that would be used if a parallel region started now +
-omp_set_num_threads(n)    // set the default thread count at runtime +
-omp_get_num_procs()       // number of logical processors available to the program +
-omp_get_wtime()           // wall-clock time in seconds; use for timing parallel regions +
-omp_in_parallel()         // 1 if called from inside a parallel region, 0 otherwise +
-``` +
- +
-### Environment variables +
- +
-^ Variable ^ Default ^ Description ^ +
-| `OMP_NUM_THREADS` | core count | Number of threads to use in each parallel region | +
-| `OMP_SCHEDULE` | `static` | Default schedule kind and optional chunk size, e.g. `dynamic,4` | +
-| `OMP_PROC_BIND` | `false` | Thread-to-core affinity policy: `close`, `spread`, or `master` | +
-| `OMP_PLACES` | (unset) | Placement units for affinity: `cores`, `threads`, or `sockets` | +
-| `OMP_MAX_ACTIVE_LEVELS` | `1` | Maximum nesting depth of simultaneously active parallel regions | +
-| `OMP_DISPLAY_ENV` | `false` | Print OpenMP version and active settings at startup: `TRUE` or `VERBOSE` |+
  
openmp.1781175804.md.gz · Last modified: by Ivan Janevski