Site Tools


openmp

**This is an old revision of the document!**

Table of Contents

This article is synthetic # OpenMP OpenMP is a shared-memory parallelism API for C, C++, and Fortran. It lets a single program exploit multiple CPU cores by distributing work across a team of threads that all share the same address space. This is in contrast to distributed-memory models like MPI, where each process has its own memory. OpenMP is implemented via compiler directives (#pragma omp in C/C++), a small runtime library (libomp), and a set of environment variables.

The execution model is fork-join: the program starts as a single master thread. When it hits a #pragma omp parallel block, it forks into a team of worker threads; all threads execute the block concurrently, then join back into one thread at the closing brace. The compiler inserts the thread creation, synchronization, and teardown code. The programmer only writes directives. If the compiler does not support OpenMP, it silently ignores all #pragma omp directives and the program runs serially, which is a useful property during debugging.

#pragma omp parallel
{
    int tid = omp_get_thread_num();
    int nthreads = omp_get_num_threads();
    printf("thread %d of %d\n", tid, nthreads);
}

The number of threads defaults to the number of logical CPU cores. It can be overridden with the OMP_NUM_THREADS environment variable or the num_threads(N) clause on the pragma. Output order is non-deterministic — threads are scheduled by the OS. Compile with -fopenmp and include <omp.h> to use the omp_* runtime functions.

Adding more threads does not always mean proportionally faster code. Amdahl's law states that if a fraction $s$ of the program is inherently serial, the maximum possible speedup is $1/s$ regardless of how many threads are used. A loop that accounts for 80% of runtime can at best give 5× speedup no matter how many cores are available. This makes reducing the serial fraction more impactful than simply increasing thread count. Eliminating overhead sources like false sharing, load imbalance, and unnecessary synchronisation compounds that further.

Concepts

Overview

Directives

#pragma omp parallel                         // fork a team of threads; join at closing brace
#pragma omp parallel for                     // distribute loop iterations across the team
#pragma omp parallel for reduction(+:s)     // loop with a parallel reduction
#pragma omp parallel sections                // distribute independent blocks across the team
#pragma omp section                          // one block inside a sections region
#pragma omp single                           // one thread runs the block; others wait at end
#pragma omp master                           // thread 0 only; no implicit barrier
#pragma omp task                             // package work for any idle thread to execute
#pragma omp taskwait                         // wait for all child tasks to finish
#pragma omp barrier                          // all threads wait until every thread arrives
#pragma omp critical                         // mutual exclusion — one thread at a time
#pragma omp atomic                           // single hardware-atomic read-modify-write
#pragma omp simd                             // assert the loop is safe to vectorise
#pragma omp flush                            // enforce memory visibility across threads

Functions

omp_get_thread_num()      // ID of the calling thread (0 … N-1)
omp_get_num_threads()     // number of threads in the current team
omp_get_max_threads()     // threads that would be used if a parallel region started now
omp_set_num_threads(n)    // set the default thread count at runtime
omp_get_num_procs()       // number of logical processors available to the program
omp_get_wtime()           // wall-clock time in seconds; use for timing parallel regions
omp_in_parallel()         // 1 if called from inside a parallel region, 0 otherwise

Environment variables

Variable Default Description
OMP_NUM_THREADS core count Number of threads to use in each parallel region
OMP_SCHEDULE static Default schedule kind and optional chunk size, e.g. dynamic,4
OMP_PROC_BIND false Thread-to-core affinity policy: close, spread, or master
OMP_PLACES (unset) Placement units for affinity: cores, threads, or sockets
OMP_MAX_ACTIVE_LEVELS 1 Maximum nesting depth of simultaneously active parallel regions
OMP_DISPLAY_ENV false Print OpenMP version and active settings at startup: TRUE or VERBOSE
openmp.1781170479.txt.gz · Last modified: by Ivan Janevski