Site Tools


wiki:gcc-profiling

Table of Contents

GCC Profiling

GCC profiling with -pg instruments the program to record function call counts and execution time. Use gprof to analyze the results. This is a simple, portable alternative to perf but less detailed.

gcc -pg -O3 -o app app.c
./app                   # runs normally but creates gmon.out
gprof app gmon.out      # analyze profiling data

-pg adds instrumentation to every function call. When the program runs, it records call counts and timing. The gmon.out file contains the profiling data.

Example output:

  %   cumulative   self              self     total
 time   seconds   seconds    calls  ms/call  ms/call  name
 45.0      0.45     0.45    10000     0.045    0.050  matrix_multiply
 30.0      0.75     0.30    20000     0.015    0.020  vector_dot
 25.0      1.00     0.25    30000     0.008    0.010  element_multiply

Each function's time and call count is shown. This reveals which functions consume the most CPU time. The “self” column is time spent in the function itself; “total” includes time in functions it calls.

Advantages: 1. Portable—works on any system with GCC 2. No setup—just compile with -pg and run 3. Simple output—easy to understand which functions are hot

Disadvantages: 1. Less precise than perf—gprof samples at fixed intervals, can miss small functions 2. Overhead—instrumentation adds overhead to every function call 3. No cache/memory analysis—gprof only tracks timing, not cache misses or bandwidth 4. Flat profile—doesn't show call chains as clearly as [[perf]] report

Call graph: Use gprof -A for call graph (which functions call which):

gprof -A app gmon.out

This shows the full call tree, revealing which function call chains consume the most time.

Annotation: Use gprof -S to annotate source code with per-line timing:

gprof -S app gmon.out | less

Lines where the program spent time are marked with sample counts.

Profiling multithreaded code: GCC profiling doesn't work well with multithreading (each thread overwrites gmon.out). Use perf instead for profiling threaded programs.

Profile-guided optimization (PGO): GCC's -fprofile-generate and -fprofile-use use profiling data to optimize the binary. This is different from gprof—it's automatic optimization, not manual profiling.

gcc -O3 -fprofile-generate -o app app.c
./app < training_input   # run with representative input
gcc -O3 -fprofile-use -o app app.c    # recompile using profiling data

This makes GCC's optimizations smarter by telling it which branches are hot and which functions are frequently called.

wiki/gcc-profiling.md · Last modified: by 127.0.0.1