PerfExpert - The Institute for Computational Engineering

Download Report

Transcript PerfExpert - The Institute for Computational Engineering

TeraGrid Tutorial
PerfExpert: An Automated Approach to
Analyzing and Optimizing the Node-Level
Performance of HPC Applications
James Browne and Martin Burtscher
The University of Texas at Austin
Goals for Tutorial
 Operational effectiveness with PerfExpert
 Be comfortable using PerfExpert
 See how easy it is to use
 Understanding of how PerfExpert works
 Appreciate the sophistication of the analysis engine
 Solicit collaborations
 For applications of PerfExpert
 For installation of PerfExpert at participant-local sites
 Seek feedback on all aspects
PerfExpert Tutorial
2
Large Cluster Performance Issues
 Scaling
 Inter-node communication and load balancing
 Input / Output
 Parallelization, overlap, and check-pointing
 Core, chip, and node level
 Memory bandwidth and latency
 Core/chip-level parallelization
 Compiler switch complexity
 PerfExpert currently focuses on core, chip and
node-level analyzes and optimizations
PerfExpert Tutorial
3
Tutorial Overview
 Introduction
 Why yet another performance tool?
 What PerfExpert does and how it works
 Demonstration
 Quick-start example and hands-on demo
 Optimization
 Optimization process
 Code tuning examples
 Foundations
 Porting and installation
PerfExpert Tutorial
4
Introduction
 HPC systems
 Operate at a small fraction of peak performance
 Performance optimization complexity is growing

e.g., migration to multicore-based compute nodes
 Diagnosing performance problems
 Requires detailed performance expertise
 HPC application writers are domain experts

not familiar with architectural details
PerfExpert Tutorial
5
Existing Performance Tools
 TAU, IPM, HPCToolkit, Pin, etc.
 Tools based on multiple approaches
 Performance counters
 Event traces
 Code instrumentation
 Target functionality – not ease of use
PerfExpert Tutorial
6
Problem
 Status: Most performance evaluation tools
 Effective use requires knowledge of architectural
details and compiler algorithms and switches
 Users must master complex interfaces
 Result: HPC application developers
 Do not use performance assessment tools
 Use them ineffectively
 Do not know how to apply information from tool
PerfExpert Tutorial
7
Performance Counter-based Tools
 Which tool?
 Which counter(s)?
 100s of possibilities
 Cryptic descriptions
 What is counted?
 Adds include subtractions
 L1 misses exclude lines
with prefetch requests
PerfExpert Tutorial
L1_DCM
L1_ICM
L2_DCM
L2_ICM
L2_TCM
TLB_DM
TLB_IM
BR_TKN
BR_MSP
TOT_INS
FP_INS
BR_INS
VEC_INS
TOT_CYC
L1_DCA
L2_DCA
L2_ICH
L1_ICA
L2_ICA
L1_ICR
L2_TCA
FML_INS
FAD_INS
FDV_INS
FSQ_INS
FP_OPS
...
Level 1 data cache misses
Level 1 instruction cache misses
Level 2 data cache misses
Level 2 instruction cache misses
Level 2 cache misses
Data translation lookaside buffer misses
Instruction translation lookaside buffer misses
Conditional branch instructions taken
Conditional branch mispredictions
Instructions completed
Floating point instructions
Branch instructions
Vector/SIMD instructions
Total cycles
Level 1 data cache accesses
Level 2 data cache accesses
Level 2 instruction cache hits
Level 1 instruction cache accesses
Level 2 instruction cache accesses
Level 1 instruction cache reads
Level 2 total cache accesses
Floating point multiply instructions
Floating point add instructions
Floating point divide instructions
Floating point square root instructions
Floating point operations
8
PerfExpert Project Goal
 Automate detection of performance bottlenecks
 At core, chip, and node level
 Suggest optimizations for each bottleneck
 Including code examples and compiler switches
 Later: apply suggestions automatically
 Simplicity is paramount
 Trivial user interface
 Easily understandable output
PerfExpert Tutorial
9
Related Work
 Automatic bottleneck analysis and remediation
 PERCS project at IBM Research
Less automation for bottleneck identification and analysis
 Not open source

 PERI Autotuning project
 Parallel Performance Wizard

Event trace analysis, program instrumentation
 Analysis tools with automated diagnosis
 Projects that target multicore optimizations
PerfExpert Tutorial
10
PerfExpert v1.1
 Features
 Semi-automatic bottleneck detection (uses HPCToolkit)
 Extensive list of suggested optimizations with examples
 Simple and intuitive interface
 Capabilities
 Based on PAPI and native performance counters
 Intra-node performance analysis
 Installed on Ranger and Longhorn
PerfExpert Tutorial
11
Workflow Simplification
Typical optimization workflow with profiling tools
Optimization workflow with
PerfExpert
[mostly manual]
[mostly automated]
Selecting performance counters
Running multiple
measurements
Collecting
performance data
Identifying
bottlenecks
Automatic (for
core, chip, & nodelevel bottlenecks)
performance
counter selection,
measurement
execution,
data collection,
bottleneck diagnosis,
and optimization
suggestion based
on several categories
Searching for proper
optimization method
Implementing
optimization
PerfExpert Tutorial
Selecting and implementing optimization
12
PerfExpert Approach
 Gather performance counter measurements
 Multiple runs with HPCToolkit
 Sampling-based results for procedures and loops
 Combine results
 Check variability, runtime, consistency, and integrity
 Compute and output assessment
 Only for most important code sections
 Correlate results from different runs
PerfExpert Tutorial
13
PerfExpert Performance Metric
 Local Cycles Per Instruction (LCPI)
 Compute upper bounds on CPI contribution for various
categories, e.g., branches and memory accesses
(BR_INS * BR_lat + BR_MSP * BR_miss_lat) / TOT_INS
 (L1_DCA * L1_dlat + L2_DCA * L2_lat + L2_DCM * Mem_lat) / TOT_INS

 Benefits
 Highlights key aspects and hides misleading details
 Relative metric (less susceptible to non-determinism)
 Easily extensible to include additional groups
 Can be refined with better or more perf counters
PerfExpert Tutorial
14
PerfExpert Output
PerfExpert v1.1 (Ranger)
Copyright (c) 2010, The University of Texas at Austin.
usage: PerfExpert threshold file [file]
total runtime in experiment.xml is 0.03 seconds
matrixproduct (53.5% of the total runtime)
---------------------------------------------------------------------------WARNING: The runtime is too short to gather meaningful measurements.
loop at line 25 in matrixproduct (22.6% of the total runtime)
---------------------------------------------------------------------------WARNING: The cycle count variation is 33.3%, making the results unreliable.
WARNING: The runtime is too short to gather meaningful measurements.
PerfExpert Tutorial
15
PerfExpert Output for MMM
total runtime in mmm-database-1234567/experiment.xml is 3.74 seconds
Suggestions on how to alleviate performance bottlenecks are available at:
http://www.tacc.utexas.edu/perfexpert/
main (100.0% of the total runtime)
----------------------------------------------------------------------------performance assessment
LCPI good......okay......fair......poor......bad....
- overall
9.6 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
upper bound by category
- data accesses
14.7 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
- instruction accesses
0.6 >>>>>>
- data TLB
9.9 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
- instruction TLB
0.0 >
- branch instructions
0.1 >
- floating-point instr
3.0 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
PerfExpert Tutorial
16
PerfExpert Output for Mangll
total runtime in mangll_dgae_snell_N3_4.xml is 193.70 seconds
total runtime in mangll_dgae_snell_N3_16.xml is 76.67 seconds
Suggestions on how to alleviate performance bottlenecks are available at:
http://www.tacc.utexas.edu/perfexpert/
dgae_RHS (runtimes are 135.86s and 45.31s)
----------------------------------------------------------------------------performance assessment
LCPI good......okay......fair......poor......bad....
- overall
>>>>>>>>>>>2222
upper bound by category
- data accesses
>>>>>>>>>>>>>>>>>>>>>>>>>>
- instruction accesses
>>>>
- data TLB
>
- instruction TLB
>
- branch instructions
>
- floating-point instr
>>>>>>>>>>>>>>>>>>>>>
PerfExpert Tutorial
17
Suggestions with Examples
If floating-point instructions are a problem
 Reduce the number of floating-point instructions
a) eliminate floating-point operations through distributivity
d[i] = a[i] * b[i] + a[i] * c[i]; → d[i] = a[i] * (b[i] + c[i]);
 Avoid divides
b) compute the reciprocal once outside of loop and use multiplication inside the loop
loop i {a[i] = b[i] / c;} → cinv = 1.0 / c; loop i {a[i] = b[i] * cinv;}
 Avoid square roots
c) compare squared values instead of computing the square root
if (x < sqrt(y)) {} → if ((x < 0.0) || (x*x < y)) {}
 Speed up divide and square-root operations
d) use float instead of double data type if loss of precision is acceptable
double a[n]; → float a[n];
e) allow the compiler to trade off precision for speed
try the “-prec-div”, “-prec-sqrt”, and “-pc32” compiler flags
PerfExpert Tutorial
18
Mangll Optimization Case Study
If data accesses are a problem
 Reduce the number of memory accesses
a) copy data into local scalar variables and operate on the local copies
b) recompute values rather than loading them if doable with few operations
c) vectorize the code
 Improve the data locality
d) componentize important loops by factoring them into their own subroutines
e) employ loop blocking and interchange (change the order of the memory accesses)
f) reduce the number of memory areas (e.g., arrays) accessed simultaneously
g) split structs into hot and cold parts, where the hot part has a pointer to the cold part
 Other
h) use smaller types (e.g., float instead of double or short instead of int)
i) for small elements, allocate an array of elements instead of each element individually
j) align data, especially arrays and structs
k) pad memory areas so that temporal elements do not map to the same set in the cache
PerfExpert Tutorial
19
Eliminate Inapplicable Suggestions
If data accesses are a problem
 Reduce the number of memory accesses
a) copy data into local scalar variables and operate on the local copies
b) recompute values rather than loading them if doable with few operations
c) vectorize the code
 Improve the data locality
d) componentize important loops by factoring them into their own subroutines
e) employ loop blocking and interchange (change the order of the memory accesses)
f) reduce the number of memory areas (e.g., arrays) accessed simultaneously
g) split structs into hot and cold parts, where the hot part has a pointer to the cold part
 Other
h) use smaller types (e.g., float instead of double or short instead of int)
i) for small elements, allocate an array of elements instead of each element individually
j) align data, especially arrays and structs
k) pad memory areas so that temporal elements do not map to the same set in the cache
PerfExpert Tutorial
20
Try Remaining Suggestions
If data accesses are a problem
 Reduce the number of memory accesses
a) copy data into local scalar variables and operate on the local copies
b) recompute values rather than loading them if doable with few operations
c) vectorize the code
 Improve the data locality
d) componentize important loops by factoring them into their own subroutines
e) employ loop blocking and interchange (change the order of the memory accesses)
f) reduce the number of memory areas (e.g., arrays) accessed simultaneously
g) split structs into hot and cold parts, where the hot part has a pointer to the cold part
 Other
h) use smaller types (e.g., float instead of double or short instead of int)
i) for small elements, allocate an array of elements instead of each element individually
j) align data, especially arrays and structs
k) pad memory areas so that temporal elements do not map to the same set in the cache
PerfExpert Tutorial
21
Next: Demonstration
 Step-by-step demonstration of PerfExpert usage
 Simple matrix-matrix multiplication example
 Live on Ranger if possible
PerfExpert Tutorial
22
Tutorial Overview
 Introduction
 Why yet another performance tool?
 What PerfExpert does and how it works
 Demonstration
 Quick-start example and hands-on demo
 Optimization
 Optimization process
 Code tuning examples
 Foundations
 Porting and installation
PerfExpert Tutorial
23
Approach
 Run quick-start example: matrix-matrix multiply
 http://www.tacc.utexas.edu/perfexpert/

Quick-start guide [pdf] [html]
 Content
Assumes program already runs on Ranger
 First page explains how to run PerfExpert
 Second page explains output format
 Third page walks through optimization example

 Try MMM example and/or your own program
PerfExpert Tutorial
24
Quick-Start Guide Example
 Matrix-matrix multiply
 Deliberately miscoded for simplicity of illustration
for (i = 0; i < n; i++)
for (j = 0; j < n; j++)
for (k = 0; k < n; k++)
c[i][j] += a[i][k] * b[k][j];
 What is wrong with this code?
 PerfExpert will report WHY this code performs
poorly and suggest optimizations
PerfExpert Tutorial
25
Step 1: Measure Application (1)
 Measure performance with following commands
module load papi java perfexpert
cp $TACC_PERFEXPERT_DIR/PerfExpert.sge ./
vi PerfExpert.sge
qsub PerfExpert.sge
 Penultimate command starts a text editor
PerfExpert Tutorial
26
Measure Application (2)
 Adjust highlighted entries in submission script
#!/bin/bash
#$ -A RangerTechInsertion
#$ -V
#$ -cwd
#$ -N PerfExpert
#$ -j y
#$ -o $JOB_NAME.$JOB_ID.out
#$ -pe 1way 16
#$ -q development
#$ -l h_rt=00:07:00
export MyExePath=./
export MyExeName=a.out
export MyCmdLine="myargs"
...
PerfExpert Tutorial
#
#
#
#
#
#
#
#
#
project name (not for guest accounts)
inherit the submission environment
start job in submission directory
job name
combine stderr & stdout into stdout
name of the output file
requests x cores/node, y cores total
queue name
7 times the expected runtime (hh:mm:ss)
# path to application
# application name (must not contain path)
# command line arguments for application
27
Step 2: Determine Bottlenecks
 Identify bottlenecks with the following command
 Command also listed at bottom of job output
PerfExpert 0.1 ./hpctoolkit-a.out-database1234567/experiment.xml
 Adjust highlighted entries
 Threshold (0.1)

Output code sections representing ≥ 0.1 times total runtime
 Application name (a.out)
 Ranger job ID (1234567)
PerfExpert Tutorial
28
Output for Sample MMM Code
total runtime in hpctoolkit-a.out-database-1234567/experiment.xml is 3.74 sec
Suggestions on how to alleviate performance bottlenecks are available at:
http://www.tacc.utexas.edu/perfexpert/
URL to suggested optimizations
...
procedure or loop identifier (if compiled with “-g”)
loop at line 25 in main (99.7% of the total runtime)
overall loop performance is bad
----------------------------------------------------------------------------performance assessment
LCPI good......okay......fair......poor......bad....
- overall
9.6 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
upper bound by category
- data accesses
14.7 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
- instruction accesses
0.6 >>>>>>
- data TLB
9.9 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
- instruction TLB
0.0 >
most of runtime due to data
TLB and data accesses
- branch instructions
0.1 >
- floating-point instr
3.0 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
PerfExpert Tutorial
29
Step 3: Optimize Critical Code Section
 Loop nest at line 25
for (i = 0; i < n; i++)
for (j = 0; j < n; j++)
for (k = 0; k < n; k++)
c[i][j] += a[i][k] * b[k][j];
 Identified main bottlenecks
 Memory accesses
 Data TLB
 Focus on data TLB problem first (for simplicity)
PerfExpert Tutorial
30
Data TLB Optimization Suggestions
1) Improve the data locality
a) use superpages (larger page sizes)
not yet enabled on all Ranger nodes
b) change the order of loops
loop i {...} loop j {...} → loop j {...} loop i {...}
c) employ loop blocking and interchange (change the order of the memory accesses)
loop i {loop k {loop j {c[i][j] = c[i][j] + a[i][k] * b[k][j];}}} →
loop k step s {loop j step s {loop i {for (kk = k; kk < k + s; kk++) {for (jj = j; jj < j + s; jj++)
{c[i][jj] = c[i][jj] + a[i][kk] * b[kk][jj];}}}}}
2) Reduce the data size
a) use smaller types (e.g., float instead of double or short instead of int)
double a[n]; → float a[n];
use the "-fpack-struct" compiler flag
b) allocate an array of elements instead of each element individually
loop {... c = malloc(1); ...} → top = n; loop {if (top == n) {tmp = malloc(n); top = 0;} ... c =
&tmp[top++]; ...}
PerfExpert Tutorial
31
Eliminate Inapplicable Suggestions
1) Improve the data locality
a) use superpages (larger page sizes)
not yet enabled on all Ranger nodes
b) change the order of loops
loop i {...} loop j {...} → loop j {...} loop i {...}
c) employ loop blocking and interchange (change the order of the memory accesses)
loop i {loop k {loop j {c[i][j] = c[i][j] + a[i][k] * b[k][j];}}} →
loop k step s {loop j step s {loop i {for (kk = k; kk < k + s; kk++) {for (jj = j; jj < j + s; jj++)
{ c[i][jj] = c[i][jj] + a[i][kk] * b[kk][jj];}}}}}
2) Reduce the data size
a) use smaller types (e.g., float instead of double or short instead of int)
double a[n]; → float a[n];
use the "-fpack-struct" compiler flag
b) allocate an array of elements instead of each element individually
loop {... c = malloc(1); ...} → top = n; loop {if (top == n) {tmp = malloc(n); top = 0;} ... c =
&tmp[top++]; ...}
PerfExpert Tutorial
32
Try Remaining Suggestions
 Start with suggestion 1b because it is simpler
1) Improve the data locality
b) change the order of loops
loop i {...} loop j {...} → loop j {...} loop i {...}
 Exchange the j and k loops of the loop nest
for (i = 0; i < n; i++)
for (k = 0; k < n; k++)
for (j = 0; j < n; j++)
c[i][j] += a[i][k] * b[k][j];
 Assess new code with PerfExpert
PerfExpert Tutorial
33
Output after Loop Exchange
total runtime in hpctoolkit-a.out-database-1234568/experiment.xml is 1.45 sec
runtime is much lower
Suggestions on how to alleviate performance bottlenecks are available at:
http://www.tacc.utexas.edu/perfexpert/
...
overall loop performance is better but still bad
loop at line 25 in main (100.0% of the total runtime)
----------------------------------------------------------------------------performance assessment
LCPI good......okay......fair......poor......bad....
- overall
4.1 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
upper bound by category
- data accesses
3.6 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
- instruction accesses
0.5 >>>>>
data accesses should be optimized next
- data TLB
0.0 >
- instruction TLB
0.0 >
data TLB is no longer a problem
- branch instructions
0.1 >
- floating-point instr
3.3 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
PerfExpert Tutorial
34
Continuing Optimization
 Performance is much improved
 Runtime dropped to less than half
 Data TLB problem is fixed
 PerfExpert correctly identified this bottleneck
 Suggested a useful code optimization
 Helped verify the resolution of the problem
 Performance is still not good
 Repeat optimization procedure
 Target data accesses next (e.g., loop blocking)
PerfExpert Tutorial
35
Live Demo and Do-It-Yourself
 Demo of quick-start example
 Obtain guest account information
 Log on to Ranger
 Copy the MMM code: cp ~burtsche/PET1/mmm* ./
 Follow along with the demo
 Try your own program
 Prepare binary
 Follow steps in quick-start guide (see also next slide)
PerfExpert Tutorial
36
Try Your Own Program
 Log on to your account
 Load the Java, PAPI, and PerfExpert modules
 module load java papi perfexpert
 Compile your code with optimization and “-g”
 mpicc -O3 -g source.c
 mpif90 -O3 -g source.f90
 Copy the PerfExpert.sge submission script
 cp $TACC_PERFEXPERT_DIR/PerfExpert.sge ./
 Edit the PerfExpert.sge script appropriately
PerfExpert Tutorial
37
Conclusion
 PerfExpert automates measurement and analysis
 Recommends optimizations
 Optimization application is still manual
 User interface is simple
 No complicated command line or configuration files
 Output is easy to understand
 Code sections sorted by importance
 The longer the bar, the more important to optimize
 Next: approaches and examples for optimization
PerfExpert Tutorial
38
Overview
 Introduction
 Why yet another performance tool?
 What PerfExpert does and how it works
 Demonstration
 Quick-start example and hands-on demo
 Optimization
 Optimization process
 Code tuning examples
 Foundations
 Porting and installation
PerfExpert Tutorial
39
Optimization Overview
 Optimization process
 Matrix-matrix multiplication example
 TLB optimization
 Data access optimization
 Heat-transfer code example
 Summary and future of optimizations
PerfExpert Tutorial
40
Optimization – Pattern Recognition
 Ideal – Compilers generate optimized code for
every possible code construct
 Reality – Compilers generate near optimal code
for a few patterns for each general category of
source code constructs
 Optimization
 Determine code segment local performance
bottlenecks (what compiler needs to optimize)
 Write (or refactor) code to patterns for which
compilers will generate near optimal code
PerfExpert Tutorial
41
Determining What to Optimize
 PerfExpert reports “cost” of
 Data accesses
 Instruction accesses
 Data TLB
 Instruction TLB
 Branch Instructions
 Floating-point instructions
for each key code segment
PerfExpert Tutorial
42
PerfExpert Optimization Suggestions
 For each category of LCPI report
 For each possible cause for a bottleneck

Code templates that may enable compiler to optimize code
 For the MMM example:
If the data TLB is a problem
1) Improve the data locality
a) use superpages (larger page sizes)
not yet enabled on all Ranger nodes
b) change the order of loops
loop i {...} loop j {...} → loop j {...} loop i {...}
c) . . .
2) Reduce the data size
a) . . .
PerfExpert Tutorial
43
Optimization Process
1. Look at suggestions for “worst” category
 E.g., for MMM – data accesses and data TLB
2. Examine existing code for most probable cause
(use knowledge of language)
3. Examine suggestions for closest match to code
structure
4. Restructure code to optimization template
5. Rerun PerfExpert
6. Repeat from Step 1
PerfExpert Tutorial
44
Optimization Example – MMM(1,2)
1. Look at suggestions for “worst” category
 E.g., for MMM – data accesses and data TLB
2. Examine existing code for most probable cause
for (i = 0; i < n; i++)
for (j = 0; j < n; j++)
for (k = 0; k < n; k++)
c[i][j] += a[i][k] * b[k][j];
Poor data locality – From language knowledge, for multidimensional arrays, access is faster if you iterate on the array
subscript offering the smallest stride or step size. In C programs,
this is the rightmost subscript due to row-major memory layout.
PerfExpert Tutorial
45
Optimization Example – MMM(3,4)
3. Examine suggestions for closest (and simplest)
match to code structure
b) change the order of loops
loop i {...} loop j {...} → loop j {...} loop i {...}
4. Restructure code to optimization template
for (i = 0; i < n; i++)
for (k = 0; k < n; k++)
for (j = 0; j < n; j++)
c[i][j] += a[i][k] * b[k][j];
PerfExpert Tutorial
46
Optimization Example – MMM(5)
total runtime in hpctoolkit-a.out-database-1234568/experiment.xml is 1.45 sec
runtime is much lower
Suggestions on how to alleviate performance bottlenecks are available at:
http://www.tacc.utexas.edu/perfexpert/
...
overall loop performance is better but still bad
loop at line 25 in main (100.0% of the total runtime)
----------------------------------------------------------------------------performance assessment
LCPI good......okay......fair......poor......bad....
- overall
4.1 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
upper bound by category
- data accesses
3.6 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
- instruction accesses
0.5 >>>>>
data accesses should be optimized next
- data TLB
0.0 >
- instruction TLB
0.0 >
data TLB is no longer a problem
- branch instructions
0.1 >
- floating-point instr
3.3 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
PerfExpert Tutorial
47
MMM Example – Data Access(1,2)
If data accesses are a problem
1)
2)
3)
4)
5)
6)
7)
Reduce the number of memory accesses
Improve the data locality
Reduce the data size
Reduce cache line boundary crossings
Reduce conflict misses
Increase memory bandwidth
Reduce DRAM page contention
PerfExpert Tutorial
48
MMM Example – Data Access(3)
c) employ loop blocking and interchange
loop i {loop k {loop j {
c[i][j] = c[i][j] + a[i][k] * b[k][j];
}}}→
loop k step s {loop j step s {loop i {
for (kk = k; kk < k+s; kk++) {for (jj = j; jj < j+s; jj++) {
c[i][jj] = c[i][jj] + a[i][kk] * b[kk][jj];
}}
}}}
 Why?
PerfExpert Tutorial
49
MMM Example – Data Access(4)
 Blocked loop code (blocking factor s = 70)
for (k = 0; k < n; k += s) {
for (j = 0; j < n; j += s) {
for (i = 0; i < n; i++) {
for (kk = k; kk < k + s; kk++) {
for (jj = j; jj < j + s; jj++) {
c[i][jj] += a[i][kk] * b[kk][jj];
}
}
}
}
}
PerfExpert Tutorial
50
MMM Example – Data Access(5)
total runtime in ./hpctoolkit-mmm3-database-1495284/experiment.xml is 0.28 s
Suggestions on how to alleviate performance bottlenecks are available at:
http://www.tacc.utexas.edu/perfexpert/
...
loop at line 28 in main (98.8% of the total runtime)
----------------------------------------------------------------------------performance assessment
LCPI good......okay......fair......poor......bad....
- overall
0.6 >>>>>>
upper bound by category
- data accesses
2.1 >>>>>>>>>>>>>>>>>>>>>
- instruction accesses
0.6 >>>>>>
- data TLB
0.0 >
- instruction TLB
0.0 >
- branch instructions
0.2 >>
- floating-point instr
2.5 >>>>>>>>>>>>>>>>>>>>>>>>>
PerfExpert Tutorial
51
Heat Transfer Code Example
 Code Structure
 Execution Example
 PerfExpert Run
 Optimization
PerfExpert Tutorial
52
Flow Chart
Heat Conduction Model
Input
Define Array
Grid generation
Initialize
Update
Make Matrix
TDMA
BCT
BCT
No
iter > max_iter
Yes
Plot
No
time > max_iter_time
PerfExpert Tutorial
Yes
53
Heat Transfer Code
 Cartesian coordinate index extension for 3D mesh
 Extra works in main.f90 and exchange.f90
 Virtual (cartesian) topology implementation
 MPI data type setup in main.f90
 2-way communication routines in exchange.f90
 Other processes are the same as the 2-D code
PerfExpert Tutorial
54
Output of Original Code
## CASE #1 K-I-J Loop ##
========================================================
Number of MPI processes is : 32
========================================================
# of iteration / comp. time / comm. time / total runtime
========================================================
100 10.08139 3.31155 10.39492
200 20.14803 6.61543 20.77528
300 30.20992 9.91811 31.15046
400 40.27607 13.22588 41.52979
500 50.33679 16.52017 51.90433
600 60.39856 19.82516 62.28027
700 70.46096 23.12081 72.65577
800 80.52318 26.42411 83.03120
900 90.58822 29.73077 93.40942
1000 100.64964 33.02405 103.78358
PerfExpert Tutorial
55
PerfExpert Output on Original Code
total runtime in heat3d-mpi-kij.exe is 100.95 seconds
Suggestions on how to alleviate performance bottlenecks are available at:
http://www.tacc.utexas.edu/perfexpert/
tdma (65.6% of the total runtime)
----------------------------------------------------------------------performance assessment
LCPI good......okay......fair......poor......bad....
- overall
6.5 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
upper bound by category
- data accesses
14.9 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
- instruction accesses
0.4 >>>>
- data TLB
0.0 >
- instruction TLB
0.0 >
- branch instructions
0.1 >
- floating-point instr
1.2 >>>>>>>>>>>>
loop at line 54 in tdma (52.1% of the total runtime)
PerfExpert Tutorial
56
Loop Structure in Original Code
do K = SZ1, EZ1
do I = SX1, EX1
do J = SY1, EY1
D(I,J,K)=Sterm(I,J,K)+A_E(I,J,K)*PHI_OLD(I+1,J,K) &
+A_W(I,J,K)*PHI_OLD(I-1,J,K) &
+A_T(I,J,K)*PHI_OLD(I,J,K+1) &
+A_B(I,J,K)*PHI_OLD(I,J,K-1)
TDMA_P(I,J,K)=A_N(I,J,K)/(A_P(I,J,K)A_S(I,J,K)*TDMA_P(I,J-1,K))
TDMA_Q(I,J,K)=(D(I,J,K)+A_S(I,J,K)*TDMA_Q(I,J1,K))/(A_P(I,J,K)A_S(I,J,K)*TDMA_P(I,J-1,K))
enddo
enddo
PerfExpert Tutorial
57
What is Wrong with Loop Structure?
 For multi-dimensioned arrays, access is faster if
you iterate on the array subscript offering the
smallest stride or step size. In Fortran programs,
this is the leftmost subscript due to its columnmajor memory access pattern.
 PerfExpert suggestion
 Loop blocking or loop interchange
 Interchange is simpler – let’s try it first
PerfExpert Tutorial
58
Code after Loop Index Interchange
do K = SZ1, EZ1
do J = SY1, EY1
do I = SX1, EX1
D(I,J,K)=Sterm(I,J,K)+A_E(I,J,K)*PHI_OLD(I+1,J,K) &
+A_W(I,J,K)*PHI_OLD(I-1,J,K) &
+A_T(I,J,K)*PHI_OLD(I,J,K+1) &
+A_B(I,J,K)*PHI_OLD(I,J,K-1)
TDMA_P(I,J,K)=A_N(I,J,K)/(A_P(I,J,K)-A_S(I,J,K)*
TDMA_P(I,J-1,K))
TDMA_Q(I,J,K)=(D(I,J,K)+A_S(I,J,K)*TDMA_Q(I,J-1,K))/
(A_P(I,J,K)- A_S(I,J,K)*TDMA_P(I,J-1,K))
enddo
enddo
PerfExpert Tutorial
59
PerfExpert on Interchanged Loop
total runtime in heat3d-mpi-kji.exe is 76.93 seconds
Suggestions on how to alleviate performance bottlenecks are available at:
http://www.tacc.utexas.edu/perfexpert/
tdma (73.4% of the total runtime)
----------------------------------------------------------------------------performance assessment
LCPI good......okay......fair......poor......bad....
- overall
6.2 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
upper bound by category
- data accesses
5.4 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
- instruction accesses
0.4 >>>>
- data TLB
0.0 >
- instruction TLB
0.0 >
- branch instructions
0.1 >
- floating-point instr
1.6 >>>>>>>>>>>>>>>>
loop at line 53 in tdma (54.7% of the total runtime)
PerfExpert Tutorial
60
Summary and Future of Optimization
 Optimization mostly still requires human pattern




recognition skills
Experience is the best guide to pattern recognition
PerfExpert will add optimization case studies as
they are generated
We hope you will send us case studies based on
your own codes and applications of PerfExpert
We hope to automate certain common
optimizations
PerfExpert Tutorial
61
Overview
 Introduction
 Why yet another performance tool?
 What PerfExpert does and how it works
 Demonstration
 Quick-start example and hands-on demo
 Optimization
 Optimization process
 Code tuning examples
 Foundations
 Porting and installation
PerfExpert Tutorial
62
Porting and Installing Outline
 System requirements
 Software stack
 PerfExpert configuration
 Establish categories, define LCPI formulae, select
performance counters, measure system parameters
 PerfExpert installation
 Example: AMD Barcelona configuration (Ranger)
 Demo: porting to Intel Nehalem (Longhorn)
PerfExpert Tutorial
63
System Requirements for PerfExpert
 Operating system
 Linux
 Software
 Next slide
 Compilers
 C/C++ (for building tools and libraries)
 Architecture
 x86_64 (currently AMD Barcelona and Intel Nehalem)
PerfExpert Tutorial
64
Software Stack for PerfExpert
 PerfCtr patch
 http://user.it.uu.se/~mikpe/linux/perfctr/
 PAPI library
 http://icl.cs.utk.edu/papi/
 HPCToolkit
 http://hpctoolkit.org/
 Java and Perl
 http://www.java.com/
 http://www.perl.org/
PerfExpert Tutorial
65
PerfExpert Configuration
1. Establish desired categories
 Each bar in the output corresponds to a category
2. Define LCPI formula for each category
3. Select hardware performance counters
4. Determine system parameters
 Measure the necessary system properties
The categories, formulae, counters, and parameters
that follow are appropriate for Ranger
PerfExpert Tutorial
66
1. Establish Desired Categories
 Each bar in the output corresponds to a category
 Currently, the following categories are used
 Overall LCPI
 Upper LCPI bounds






Data accesses
Instruction accesses
Data TLB
Instruction TLB
Branch instructions
Floating-point instructions
PerfExpert Tutorial
67
2. Define LCPI Formulae
 Overall LCPI

TOT_CYC / TOT_INS
 Upper LCPI bounds by category






(L1_DCA * L1_dlat + L2_DCA * L2_lat + L2_DCM * Mem_lat) / TOT_INS
(L1_ICA * L1_ilat + L2_ICA * L2_lat + L2_ICM * Mem_lat) / TOT_INS
(TLB_DM * TLB_lat) / TOT_INS
(TLB_IM * TLB_lat) / TOT_INS
(BR_INS * BR_lat + BR_MSP * BR_miss_lat) / TOT_INS
((FML_INS + FAD_INS) * FP_lat + FDV_INS * FP_slow_lat) / TOT_INS
PerfExpert Tutorial
68
3. Select Hardware Counters
 Determine available performance counters
 papi_avail and papi_native_avail
 Select counters and test with microbenchmarks
 CPU cycles (TOT_CYC)
 Data TLB misses (TLB_DM)
 Instructions committed (TOT_INS)  Instruction TLB misses (TLB_IM)
 L1 data cache accesses (L1_DCA)  Branch instructions (BR_INS)
 L2 cache data accesses (L2_DCA)  Branch mispredictions (BR_MSP)
 L2 cache data misses (L2_DCM)
 Floating-point add/sub (FAD_INS)
 L1 instr cache accesses (L1_ICA)
 Floating-point mul (FML_INS)
 L2 cache instr accesses (L2_ICA)
 Floating-point div/sqrt (FDV_INS)
 L2 cache instr misses (L2_ICM)
PerfExpert Tutorial
69
4. Determine System Parameters
 For parameters with little or no variability, look
up or measure constant/worst-case value
 CPU frequency
 FP add/sub/mul latency (FP_lat)
 L1 data cache latency (L1_dlat)  FP div/sqrt latency (FP_slow_lat)
 L1 instr cache latency (L1_ilat)
 Branch instruction latency (BR_lat)
 L2 cache latency (L2_lat)
 Branch penalty (BR_miss_lat)
 For parameters with large range, use
conservative value (may have to be tuned)
 TLB miss latency (TLB_lat)
 Memory access latency (Mem_lat)
 Good CPI threshold
PerfExpert Tutorial
70
PerfExpert on Ranger
 Demo: Ranger configuration
 Demo: PerfExpert installation
PerfExpert Tutorial
71
Porting PerfExpert to Intel Nehalem
 Software stack is already installed
 Demonstration
 PerfExpert configuration steps
Same categories as on Ranger
 Performance counter selection
 System parameters
 LCPI formulae

 PerfExpert file generation

PerfExpert.sge , PerfExpert.perl, and README
PerfExpert Tutorial
72
Porting from Ranger to Longhorn
1. Establish desired categories
 No change
2. Define LCPI formula for each category
 Adjust data accesses, instr accesses, FP instructions
3. Select hardware performance counters
 Remove L2_DCM, FAD_INS, FML_INS, FDV_INS
 Add L2_TCM, FP_COMP_OPS_EXE, ARITH, L1I
4. Determine system parameters
 Same parameters (minus one), mostly different values
PerfExpert Tutorial
73
Longhorn LCPI Formulae
 Overall LCPI

TOT_CYC / TOT_INS
 Upper LCPI bounds by category

(L1_DCA * L1_dlat + L2_DCA * L2_lat + (L2_TCM - L2_ICM)* Mem_lat) / TOT_INS

(L1_ICA * L1_ilat + L1I + L2_ICA * L2_lat + L2_ICM * Mem_lat) / TOT_INS
(TLB_DM * TLB_lat) / TOT_INS
(TLB_IM * TLB_lat) / TOT_INS
(BR_INS * BR_lat + BR_MSP * BR_miss_lat) / TOT_INS
(FP_COMP_OPS_EXE * FP_lat + ARITH) / TOT_INS [no FP_slow_lat]




PerfExpert Tutorial
74
PerfExpert on Longhorn
 Demo: Ranger configuration
 Demo: PerfExpert installation
 Future work
PerfExpert Tutorial
75
Future Work (1)
 More case studies
 Applications with various bottlenecks to harden tool
 Improve and expand capabilities
 Finer-grained recommendations
 Add data structure based analyses and optimizations
 Non-performance-counter-based measurements
 Grow optimization and example database
 Automatic implementation of solutions to common
core, chip and node-level performance bottlenecks
PerfExpert Tutorial
76
Future Work (2)
 Port and deploy PerfExpert on other systems
 New generation AMD and Intel chips
 PowerPC chips
 We solicit collaborations for other chip/node
architectures
 Long term
 Heterogeneous architectures
 Communication and I/O bottleneck analyses and
optimizations
PerfExpert Tutorial
77
Goals for Tutorial
 Operational effectiveness with PerfExpert
 Be comfortable using PerfExpert
 See how easy it is to use
 Understanding of how PerfExpert works
 Appreciate the sophistication of the analysis engine
 Solicit collaborations
 For applications of PerfExpert
 For installation of PerfExpert at participant-local sites
 Seek feedback on all aspects
PerfExpert Tutorial
78