Site Tools


perf

perf

What is perf?

perf is a profiling and performance analysis tool built into the Linux kernel. It talks directly to the CPU's hardware performance monitoring unit (PMU) — a set of on-chip counters that can track things like retired instructions, cache misses, branch mispredictions, and stalled cycles — and lets you measure any of them against a running program.

Think of optimising an HPC program without a profiler. You have a tight loop that is slower than it should be, and you guess why: maybe the memory access pattern is bad, maybe the branch predictor is struggling, maybe there's false sharing between threads. You make a change, recompile, time it. The change helped — or did it? You still do not know why it was slow or whether you actually fixed the root cause. You are flying blind.

What you really want is to point at the running program and ask the hardware: how many L3 cache misses happened? How many instructions did you retire per cycle? Which function was the CPU spending most of its time in? perf answers those questions directly, because modern CPUs are instrumented for exactly this purpose. The PMU has been there the whole time; perf just gives you a way to read it.

In HPC work, perf is typically the first tool you reach for. It is low-overhead (sampling-based profiling adds roughly 1-5% overhead), requires no recompilation of the target, and the data it reports is the ground truth from the hardware rather than a model.

Install

perf is packaged as part of the Linux kernel tools. On Debian and Ubuntu:

sudo apt install linux-perf
# or, if that package is not found:
sudo apt install linux-tools-$(uname -r) linux-tools-generic

On Fedora and RHEL:

sudo dnf install perf

By default, non-root users cannot access hardware counters. Check the current restriction level:

cat /proc/sys/kernel/perf_event_paranoid

A value of 2 restricts most events to root. For development on a workstation, set it to 0:

echo 0 | sudo tee /proc/sys/kernel/perf_event_paranoid

To make this permanent across reboots, add kernel.perf_event_paranoid = 0 to /etc/sysctl.d/99-perf.conf.

For call-graph profiling (--call-graph dwarf), the target binary should be compiled with -g (DWARF debug info). Optimisation flags (-O2, -O3) are fine and should be kept on — profiling an unoptimised binary tells you nothing useful about HPC performance.

Practice

The quickest way to see perf in action is perf stat. It runs your program and prints a summary of hardware counter values when it exits:

perf stat ./my_program

The output looks like:

 Performance counter stats for './my_program':

       4,102.38 msec task-clock                #    3.98 CPUs utilized
             42      context-switches          #   10.238 /sec
              4      cpu-migrations            #    0.975 /sec
            512      page-faults               #  124.810 /sec
 16,847,203,412      cycles                    #    4.107 GHz
 12,334,901,088      instructions              #    0.73  insn per cycle
  2,301,445,221      branches                  #  561.035 M/sec
     48,223,901      branch-misses             #    2.09% of all branches
  1,203,887,654      cache-references          #  293.461 M/sec
    401,292,771      cache-misses              #   33.33% of cache refs

The two numbers to read first are instructions per cycle (IPC) and cache miss rate. IPC of 0.73 on a modern out-of-order CPU (which can retire 3-4 instructions per cycle at peak) means the pipeline is largely stalled — the CPU is waiting for something. A 33% L3 cache miss rate is a strong hint about what it is waiting for: memory. That is enough information to direct the next optimisation attempt.

Concepts

perf stat

perf stat counts hardware events over the full lifetime of a program. It is the fastest way to get a quantitative characterisation of what the CPU was doing.

perf stat -e cycles,instructions,cache-misses,L1-dcache-load-misses ./prog

The -e flag selects specific events. Without it, perf uses a default set that covers the most useful high-level metrics. You can count multiple events in one run:

perf stat -e cycles,instructions \
          -e cache-references,cache-misses \
          -e branches,branch-misses \
          ./prog

For multithreaded programs, add -a to count system-wide across all CPUs, or attach to a running process with -p <pid>.

The most important derived metric is IPC (instructions per cycle). A well-optimised HPC kernel on a modern superscalar CPU should reach IPC > 2. IPC below 1 almost always indicates memory-bound execution — the CPU is issuing loads that miss cache and must wait for DRAM.

perf record and perf report

perf stat tells you the aggregate; perf record and perf report tell you where. Record samples the program at a fixed event rate (by default, every 4000 CPU cycles) and records the current instruction pointer and call stack at each sample. After the run, the samples are aggregated by function.

perf record -g ./prog          # record with call graphs
perf report                    # open the interactive browser

perf record writes samples to perf.data in the current directory. perf report opens a text UI showing the functions where samples landed, sorted by percentage. Pressing Enter on a function expands the call graph below it.

For accurate call graphs on optimised code, use DWARF unwinding:

perf record --call-graph dwarf ./prog
perf report --stdio             # non-interactive output for scripts

perf annotate goes one level deeper: it shows the source (or disassembly) with per-instruction sample counts inline, so you can see exactly which instruction inside a hot function the CPU was stalled on.

perf annotate --stdio -s my_hot_function

Events

perf exposes three categories of events:

Hardware events come from the PMU and count micro-architectural phenomena:

cycles                  -- CPU clock cycles
instructions            -- retired instructions
cache-references        -- last-level cache accesses
cache-misses            -- last-level cache misses
branches                -- branch instructions
branch-misses           -- branch mispredictions
stalled-cycles-frontend -- pipeline stalls waiting for instruction fetch
stalled-cycles-backend  -- pipeline stalls waiting for execution resources (often: memory)

Software events come from the kernel and count OS-level phenomena: context-switches, page-faults, cpu-migrations, minor-faults, major-faults.

Tracepoints instrument specific kernel functions and subsystems: syscalls:sys_enter_read, block:block_rq_issue, sched:sched_switch. These are useful for diagnosing I/O behaviour and scheduler interaction in MPI jobs.

Use perf list to see every event available on the current machine, including CPU-vendor-specific events (Intel PEBS events, AMD IBS events) that expose finer-grained data than the standard hardware events.

perf list
perf list cache    # filter to cache-related events

Flame graphs

Flame graphs are a visualisation of perf record call-graph data developed by Brendan Gregg. Each horizontal bar represents a function; its width is proportional to the fraction of samples that included it in the call stack. Functions are stacked vertically with their callers below and callees above. The widest bars at the top of the flame are the hot leaf functions — the places where the CPU actually spent time.

perf record --call-graph dwarf -F 99 ./prog
perf script | stackcollapse-perf.pl | flamegraph.pl > flame.svg

stackcollapse-perf.pl and flamegraph.pl are from the FlameGraph repository. The resulting SVG is interactive: clicking a bar zooms into that subtree.

Flame graphs are especially useful for MPI programs where the hot path passes through several layers of library code (MPI implementation → network transport → OS). Without the full call graph, perf report shows time in a low-level transport function with no context; the flame graph shows exactly which application call chain led there.

perf.txt · Last modified: by 127.0.0.1