PerfExpert - The Institute for Computational Engineering
Download
Report
Transcript PerfExpert - The Institute for Computational Engineering
TeraGrid Tutorial
PerfExpert: An Automated Approach to
Analyzing and Optimizing the Node-Level
Performance of HPC Applications
James Browne and Martin Burtscher
The University of Texas at Austin
Goals for Tutorial
Operational effectiveness with PerfExpert
Be comfortable using PerfExpert
See how easy it is to use
Understanding of how PerfExpert works
Appreciate the sophistication of the analysis engine
Solicit collaborations
For applications of PerfExpert
For installation of PerfExpert at participant-local sites
Seek feedback on all aspects
PerfExpert Tutorial
2
Large Cluster Performance Issues
Scaling
Inter-node communication and load balancing
Input / Output
Parallelization, overlap, and check-pointing
Core, chip, and node level
Memory bandwidth and latency
Core/chip-level parallelization
Compiler switch complexity
PerfExpert currently focuses on core, chip and
node-level analyzes and optimizations
PerfExpert Tutorial
3
Tutorial Overview
Introduction
Why yet another performance tool?
What PerfExpert does and how it works
Demonstration
Quick-start example and hands-on demo
Optimization
Optimization process
Code tuning examples
Foundations
Porting and installation
PerfExpert Tutorial
4
Introduction
HPC systems
Operate at a small fraction of peak performance
Performance optimization complexity is growing
e.g., migration to multicore-based compute nodes
Diagnosing performance problems
Requires detailed performance expertise
HPC application writers are domain experts
not familiar with architectural details
PerfExpert Tutorial
5
Existing Performance Tools
TAU, IPM, HPCToolkit, Pin, etc.
Tools based on multiple approaches
Performance counters
Event traces
Code instrumentation
Target functionality – not ease of use
PerfExpert Tutorial
6
Problem
Status: Most performance evaluation tools
Effective use requires knowledge of architectural
details and compiler algorithms and switches
Users must master complex interfaces
Result: HPC application developers
Do not use performance assessment tools
Use them ineffectively
Do not know how to apply information from tool
PerfExpert Tutorial
7
Performance Counter-based Tools
Which tool?
Which counter(s)?
100s of possibilities
Cryptic descriptions
What is counted?
Adds include subtractions
L1 misses exclude lines
with prefetch requests
PerfExpert Tutorial
L1_DCM
L1_ICM
L2_DCM
L2_ICM
L2_TCM
TLB_DM
TLB_IM
BR_TKN
BR_MSP
TOT_INS
FP_INS
BR_INS
VEC_INS
TOT_CYC
L1_DCA
L2_DCA
L2_ICH
L1_ICA
L2_ICA
L1_ICR
L2_TCA
FML_INS
FAD_INS
FDV_INS
FSQ_INS
FP_OPS
...
Level 1 data cache misses
Level 1 instruction cache misses
Level 2 data cache misses
Level 2 instruction cache misses
Level 2 cache misses
Data translation lookaside buffer misses
Instruction translation lookaside buffer misses
Conditional branch instructions taken
Conditional branch mispredictions
Instructions completed
Floating point instructions
Branch instructions
Vector/SIMD instructions
Total cycles
Level 1 data cache accesses
Level 2 data cache accesses
Level 2 instruction cache hits
Level 1 instruction cache accesses
Level 2 instruction cache accesses
Level 1 instruction cache reads
Level 2 total cache accesses
Floating point multiply instructions
Floating point add instructions
Floating point divide instructions
Floating point square root instructions
Floating point operations
8
PerfExpert Project Goal
Automate detection of performance bottlenecks
At core, chip, and node level
Suggest optimizations for each bottleneck
Including code examples and compiler switches
Later: apply suggestions automatically
Simplicity is paramount
Trivial user interface
Easily understandable output
PerfExpert Tutorial
9
Related Work
Automatic bottleneck analysis and remediation
PERCS project at IBM Research
Less automation for bottleneck identification and analysis
Not open source
PERI Autotuning project
Parallel Performance Wizard
Event trace analysis, program instrumentation
Analysis tools with automated diagnosis
Projects that target multicore optimizations
PerfExpert Tutorial
10
PerfExpert v1.1
Features
Semi-automatic bottleneck detection (uses HPCToolkit)
Extensive list of suggested optimizations with examples
Simple and intuitive interface
Capabilities
Based on PAPI and native performance counters
Intra-node performance analysis
Installed on Ranger and Longhorn
PerfExpert Tutorial
11
Workflow Simplification
Typical optimization workflow with profiling tools
Optimization workflow with
PerfExpert
[mostly manual]
[mostly automated]
Selecting performance counters
Running multiple
measurements
Collecting
performance data
Identifying
bottlenecks
Automatic (for
core, chip, & nodelevel bottlenecks)
performance
counter selection,
measurement
execution,
data collection,
bottleneck diagnosis,
and optimization
suggestion based
on several categories
Searching for proper
optimization method
Implementing
optimization
PerfExpert Tutorial
Selecting and implementing optimization
12
PerfExpert Approach
Gather performance counter measurements
Multiple runs with HPCToolkit
Sampling-based results for procedures and loops
Combine results
Check variability, runtime, consistency, and integrity
Compute and output assessment
Only for most important code sections
Correlate results from different runs
PerfExpert Tutorial
13
PerfExpert Performance Metric
Local Cycles Per Instruction (LCPI)
Compute upper bounds on CPI contribution for various
categories, e.g., branches and memory accesses
(BR_INS * BR_lat + BR_MSP * BR_miss_lat) / TOT_INS
(L1_DCA * L1_dlat + L2_DCA * L2_lat + L2_DCM * Mem_lat) / TOT_INS
Benefits
Highlights key aspects and hides misleading details
Relative metric (less susceptible to non-determinism)
Easily extensible to include additional groups
Can be refined with better or more perf counters
PerfExpert Tutorial
14
PerfExpert Output
PerfExpert v1.1 (Ranger)
Copyright (c) 2010, The University of Texas at Austin.
usage: PerfExpert threshold file [file]
total runtime in experiment.xml is 0.03 seconds
matrixproduct (53.5% of the total runtime)
---------------------------------------------------------------------------WARNING: The runtime is too short to gather meaningful measurements.
loop at line 25 in matrixproduct (22.6% of the total runtime)
---------------------------------------------------------------------------WARNING: The cycle count variation is 33.3%, making the results unreliable.
WARNING: The runtime is too short to gather meaningful measurements.
PerfExpert Tutorial
15
PerfExpert Output for MMM
total runtime in mmm-database-1234567/experiment.xml is 3.74 seconds
Suggestions on how to alleviate performance bottlenecks are available at:
http://www.tacc.utexas.edu/perfexpert/
main (100.0% of the total runtime)
----------------------------------------------------------------------------performance assessment
LCPI good......okay......fair......poor......bad....
- overall
9.6 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
upper bound by category
- data accesses
14.7 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
- instruction accesses
0.6 >>>>>>
- data TLB
9.9 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
- instruction TLB
0.0 >
- branch instructions
0.1 >
- floating-point instr
3.0 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
PerfExpert Tutorial
16
PerfExpert Output for Mangll
total runtime in mangll_dgae_snell_N3_4.xml is 193.70 seconds
total runtime in mangll_dgae_snell_N3_16.xml is 76.67 seconds
Suggestions on how to alleviate performance bottlenecks are available at:
http://www.tacc.utexas.edu/perfexpert/
dgae_RHS (runtimes are 135.86s and 45.31s)
----------------------------------------------------------------------------performance assessment
LCPI good......okay......fair......poor......bad....
- overall
>>>>>>>>>>>2222
upper bound by category
- data accesses
>>>>>>>>>>>>>>>>>>>>>>>>>>
- instruction accesses
>>>>
- data TLB
>
- instruction TLB
>
- branch instructions
>
- floating-point instr
>>>>>>>>>>>>>>>>>>>>>
PerfExpert Tutorial
17
Suggestions with Examples
If floating-point instructions are a problem
Reduce the number of floating-point instructions
a) eliminate floating-point operations through distributivity
d[i] = a[i] * b[i] + a[i] * c[i]; → d[i] = a[i] * (b[i] + c[i]);
Avoid divides
b) compute the reciprocal once outside of loop and use multiplication inside the loop
loop i {a[i] = b[i] / c;} → cinv = 1.0 / c; loop i {a[i] = b[i] * cinv;}
Avoid square roots
c) compare squared values instead of computing the square root
if (x < sqrt(y)) {} → if ((x < 0.0) || (x*x < y)) {}
Speed up divide and square-root operations
d) use float instead of double data type if loss of precision is acceptable
double a[n]; → float a[n];
e) allow the compiler to trade off precision for speed
try the “-prec-div”, “-prec-sqrt”, and “-pc32” compiler flags
PerfExpert Tutorial
18
Mangll Optimization Case Study
If data accesses are a problem
Reduce the number of memory accesses
a) copy data into local scalar variables and operate on the local copies
b) recompute values rather than loading them if doable with few operations
c) vectorize the code
Improve the data locality
d) componentize important loops by factoring them into their own subroutines
e) employ loop blocking and interchange (change the order of the memory accesses)
f) reduce the number of memory areas (e.g., arrays) accessed simultaneously
g) split structs into hot and cold parts, where the hot part has a pointer to the cold part
Other
h) use smaller types (e.g., float instead of double or short instead of int)
i) for small elements, allocate an array of elements instead of each element individually
j) align data, especially arrays and structs
k) pad memory areas so that temporal elements do not map to the same set in the cache
PerfExpert Tutorial
19
Eliminate Inapplicable Suggestions
If data accesses are a problem
Reduce the number of memory accesses
a) copy data into local scalar variables and operate on the local copies
b) recompute values rather than loading them if doable with few operations
c) vectorize the code
Improve the data locality
d) componentize important loops by factoring them into their own subroutines
e) employ loop blocking and interchange (change the order of the memory accesses)
f) reduce the number of memory areas (e.g., arrays) accessed simultaneously
g) split structs into hot and cold parts, where the hot part has a pointer to the cold part
Other
h) use smaller types (e.g., float instead of double or short instead of int)
i) for small elements, allocate an array of elements instead of each element individually
j) align data, especially arrays and structs
k) pad memory areas so that temporal elements do not map to the same set in the cache
PerfExpert Tutorial
20
Try Remaining Suggestions
If data accesses are a problem
Reduce the number of memory accesses
a) copy data into local scalar variables and operate on the local copies
b) recompute values rather than loading them if doable with few operations
c) vectorize the code
Improve the data locality
d) componentize important loops by factoring them into their own subroutines
e) employ loop blocking and interchange (change the order of the memory accesses)
f) reduce the number of memory areas (e.g., arrays) accessed simultaneously
g) split structs into hot and cold parts, where the hot part has a pointer to the cold part
Other
h) use smaller types (e.g., float instead of double or short instead of int)
i) for small elements, allocate an array of elements instead of each element individually
j) align data, especially arrays and structs
k) pad memory areas so that temporal elements do not map to the same set in the cache
PerfExpert Tutorial
21
Next: Demonstration
Step-by-step demonstration of PerfExpert usage
Simple matrix-matrix multiplication example
Live on Ranger if possible
PerfExpert Tutorial
22
Tutorial Overview
Introduction
Why yet another performance tool?
What PerfExpert does and how it works
Demonstration
Quick-start example and hands-on demo
Optimization
Optimization process
Code tuning examples
Foundations
Porting and installation
PerfExpert Tutorial
23
Approach
Run quick-start example: matrix-matrix multiply
http://www.tacc.utexas.edu/perfexpert/
Quick-start guide [pdf] [html]
Content
Assumes program already runs on Ranger
First page explains how to run PerfExpert
Second page explains output format
Third page walks through optimization example
Try MMM example and/or your own program
PerfExpert Tutorial
24
Quick-Start Guide Example
Matrix-matrix multiply
Deliberately miscoded for simplicity of illustration
for (i = 0; i < n; i++)
for (j = 0; j < n; j++)
for (k = 0; k < n; k++)
c[i][j] += a[i][k] * b[k][j];
What is wrong with this code?
PerfExpert will report WHY this code performs
poorly and suggest optimizations
PerfExpert Tutorial
25
Step 1: Measure Application (1)
Measure performance with following commands
module load papi java perfexpert
cp $TACC_PERFEXPERT_DIR/PerfExpert.sge ./
vi PerfExpert.sge
qsub PerfExpert.sge
Penultimate command starts a text editor
PerfExpert Tutorial
26
Measure Application (2)
Adjust highlighted entries in submission script
#!/bin/bash
#$ -A RangerTechInsertion
#$ -V
#$ -cwd
#$ -N PerfExpert
#$ -j y
#$ -o $JOB_NAME.$JOB_ID.out
#$ -pe 1way 16
#$ -q development
#$ -l h_rt=00:07:00
export MyExePath=./
export MyExeName=a.out
export MyCmdLine="myargs"
...
PerfExpert Tutorial
#
#
#
#
#
#
#
#
#
project name (not for guest accounts)
inherit the submission environment
start job in submission directory
job name
combine stderr & stdout into stdout
name of the output file
requests x cores/node, y cores total
queue name
7 times the expected runtime (hh:mm:ss)
# path to application
# application name (must not contain path)
# command line arguments for application
27
Step 2: Determine Bottlenecks
Identify bottlenecks with the following command
Command also listed at bottom of job output
PerfExpert 0.1 ./hpctoolkit-a.out-database1234567/experiment.xml
Adjust highlighted entries
Threshold (0.1)
Output code sections representing ≥ 0.1 times total runtime
Application name (a.out)
Ranger job ID (1234567)
PerfExpert Tutorial
28
Output for Sample MMM Code
total runtime in hpctoolkit-a.out-database-1234567/experiment.xml is 3.74 sec
Suggestions on how to alleviate performance bottlenecks are available at:
http://www.tacc.utexas.edu/perfexpert/
URL to suggested optimizations
...
procedure or loop identifier (if compiled with “-g”)
loop at line 25 in main (99.7% of the total runtime)
overall loop performance is bad
----------------------------------------------------------------------------performance assessment
LCPI good......okay......fair......poor......bad....
- overall
9.6 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
upper bound by category
- data accesses
14.7 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
- instruction accesses
0.6 >>>>>>
- data TLB
9.9 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
- instruction TLB
0.0 >
most of runtime due to data
TLB and data accesses
- branch instructions
0.1 >
- floating-point instr
3.0 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
PerfExpert Tutorial
29
Step 3: Optimize Critical Code Section
Loop nest at line 25
for (i = 0; i < n; i++)
for (j = 0; j < n; j++)
for (k = 0; k < n; k++)
c[i][j] += a[i][k] * b[k][j];
Identified main bottlenecks
Memory accesses
Data TLB
Focus on data TLB problem first (for simplicity)
PerfExpert Tutorial
30
Data TLB Optimization Suggestions
1) Improve the data locality
a) use superpages (larger page sizes)
not yet enabled on all Ranger nodes
b) change the order of loops
loop i {...} loop j {...} → loop j {...} loop i {...}
c) employ loop blocking and interchange (change the order of the memory accesses)
loop i {loop k {loop j {c[i][j] = c[i][j] + a[i][k] * b[k][j];}}} →
loop k step s {loop j step s {loop i {for (kk = k; kk < k + s; kk++) {for (jj = j; jj < j + s; jj++)
{c[i][jj] = c[i][jj] + a[i][kk] * b[kk][jj];}}}}}
2) Reduce the data size
a) use smaller types (e.g., float instead of double or short instead of int)
double a[n]; → float a[n];
use the "-fpack-struct" compiler flag
b) allocate an array of elements instead of each element individually
loop {... c = malloc(1); ...} → top = n; loop {if (top == n) {tmp = malloc(n); top = 0;} ... c =
&tmp[top++]; ...}
PerfExpert Tutorial
31
Eliminate Inapplicable Suggestions
1) Improve the data locality
a) use superpages (larger page sizes)
not yet enabled on all Ranger nodes
b) change the order of loops
loop i {...} loop j {...} → loop j {...} loop i {...}
c) employ loop blocking and interchange (change the order of the memory accesses)
loop i {loop k {loop j {c[i][j] = c[i][j] + a[i][k] * b[k][j];}}} →
loop k step s {loop j step s {loop i {for (kk = k; kk < k + s; kk++) {for (jj = j; jj < j + s; jj++)
{ c[i][jj] = c[i][jj] + a[i][kk] * b[kk][jj];}}}}}
2) Reduce the data size
a) use smaller types (e.g., float instead of double or short instead of int)
double a[n]; → float a[n];
use the "-fpack-struct" compiler flag
b) allocate an array of elements instead of each element individually
loop {... c = malloc(1); ...} → top = n; loop {if (top == n) {tmp = malloc(n); top = 0;} ... c =
&tmp[top++]; ...}
PerfExpert Tutorial
32
Try Remaining Suggestions
Start with suggestion 1b because it is simpler
1) Improve the data locality
b) change the order of loops
loop i {...} loop j {...} → loop j {...} loop i {...}
Exchange the j and k loops of the loop nest
for (i = 0; i < n; i++)
for (k = 0; k < n; k++)
for (j = 0; j < n; j++)
c[i][j] += a[i][k] * b[k][j];
Assess new code with PerfExpert
PerfExpert Tutorial
33
Output after Loop Exchange
total runtime in hpctoolkit-a.out-database-1234568/experiment.xml is 1.45 sec
runtime is much lower
Suggestions on how to alleviate performance bottlenecks are available at:
http://www.tacc.utexas.edu/perfexpert/
...
overall loop performance is better but still bad
loop at line 25 in main (100.0% of the total runtime)
----------------------------------------------------------------------------performance assessment
LCPI good......okay......fair......poor......bad....
- overall
4.1 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
upper bound by category
- data accesses
3.6 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
- instruction accesses
0.5 >>>>>
data accesses should be optimized next
- data TLB
0.0 >
- instruction TLB
0.0 >
data TLB is no longer a problem
- branch instructions
0.1 >
- floating-point instr
3.3 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
PerfExpert Tutorial
34
Continuing Optimization
Performance is much improved
Runtime dropped to less than half
Data TLB problem is fixed
PerfExpert correctly identified this bottleneck
Suggested a useful code optimization
Helped verify the resolution of the problem
Performance is still not good
Repeat optimization procedure
Target data accesses next (e.g., loop blocking)
PerfExpert Tutorial
35
Live Demo and Do-It-Yourself
Demo of quick-start example
Obtain guest account information
Log on to Ranger
Copy the MMM code: cp ~burtsche/PET1/mmm* ./
Follow along with the demo
Try your own program
Prepare binary
Follow steps in quick-start guide (see also next slide)
PerfExpert Tutorial
36
Try Your Own Program
Log on to your account
Load the Java, PAPI, and PerfExpert modules
module load java papi perfexpert
Compile your code with optimization and “-g”
mpicc -O3 -g source.c
mpif90 -O3 -g source.f90
Copy the PerfExpert.sge submission script
cp $TACC_PERFEXPERT_DIR/PerfExpert.sge ./
Edit the PerfExpert.sge script appropriately
PerfExpert Tutorial
37
Conclusion
PerfExpert automates measurement and analysis
Recommends optimizations
Optimization application is still manual
User interface is simple
No complicated command line or configuration files
Output is easy to understand
Code sections sorted by importance
The longer the bar, the more important to optimize
Next: approaches and examples for optimization
PerfExpert Tutorial
38
Overview
Introduction
Why yet another performance tool?
What PerfExpert does and how it works
Demonstration
Quick-start example and hands-on demo
Optimization
Optimization process
Code tuning examples
Foundations
Porting and installation
PerfExpert Tutorial
39
Optimization Overview
Optimization process
Matrix-matrix multiplication example
TLB optimization
Data access optimization
Heat-transfer code example
Summary and future of optimizations
PerfExpert Tutorial
40
Optimization – Pattern Recognition
Ideal – Compilers generate optimized code for
every possible code construct
Reality – Compilers generate near optimal code
for a few patterns for each general category of
source code constructs
Optimization
Determine code segment local performance
bottlenecks (what compiler needs to optimize)
Write (or refactor) code to patterns for which
compilers will generate near optimal code
PerfExpert Tutorial
41
Determining What to Optimize
PerfExpert reports “cost” of
Data accesses
Instruction accesses
Data TLB
Instruction TLB
Branch Instructions
Floating-point instructions
for each key code segment
PerfExpert Tutorial
42
PerfExpert Optimization Suggestions
For each category of LCPI report
For each possible cause for a bottleneck
Code templates that may enable compiler to optimize code
For the MMM example:
If the data TLB is a problem
1) Improve the data locality
a) use superpages (larger page sizes)
not yet enabled on all Ranger nodes
b) change the order of loops
loop i {...} loop j {...} → loop j {...} loop i {...}
c) . . .
2) Reduce the data size
a) . . .
PerfExpert Tutorial
43
Optimization Process
1. Look at suggestions for “worst” category
E.g., for MMM – data accesses and data TLB
2. Examine existing code for most probable cause
(use knowledge of language)
3. Examine suggestions for closest match to code
structure
4. Restructure code to optimization template
5. Rerun PerfExpert
6. Repeat from Step 1
PerfExpert Tutorial
44
Optimization Example – MMM(1,2)
1. Look at suggestions for “worst” category
E.g., for MMM – data accesses and data TLB
2. Examine existing code for most probable cause
for (i = 0; i < n; i++)
for (j = 0; j < n; j++)
for (k = 0; k < n; k++)
c[i][j] += a[i][k] * b[k][j];
Poor data locality – From language knowledge, for multidimensional arrays, access is faster if you iterate on the array
subscript offering the smallest stride or step size. In C programs,
this is the rightmost subscript due to row-major memory layout.
PerfExpert Tutorial
45
Optimization Example – MMM(3,4)
3. Examine suggestions for closest (and simplest)
match to code structure
b) change the order of loops
loop i {...} loop j {...} → loop j {...} loop i {...}
4. Restructure code to optimization template
for (i = 0; i < n; i++)
for (k = 0; k < n; k++)
for (j = 0; j < n; j++)
c[i][j] += a[i][k] * b[k][j];
PerfExpert Tutorial
46
Optimization Example – MMM(5)
total runtime in hpctoolkit-a.out-database-1234568/experiment.xml is 1.45 sec
runtime is much lower
Suggestions on how to alleviate performance bottlenecks are available at:
http://www.tacc.utexas.edu/perfexpert/
...
overall loop performance is better but still bad
loop at line 25 in main (100.0% of the total runtime)
----------------------------------------------------------------------------performance assessment
LCPI good......okay......fair......poor......bad....
- overall
4.1 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
upper bound by category
- data accesses
3.6 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
- instruction accesses
0.5 >>>>>
data accesses should be optimized next
- data TLB
0.0 >
- instruction TLB
0.0 >
data TLB is no longer a problem
- branch instructions
0.1 >
- floating-point instr
3.3 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
PerfExpert Tutorial
47
MMM Example – Data Access(1,2)
If data accesses are a problem
1)
2)
3)
4)
5)
6)
7)
Reduce the number of memory accesses
Improve the data locality
Reduce the data size
Reduce cache line boundary crossings
Reduce conflict misses
Increase memory bandwidth
Reduce DRAM page contention
PerfExpert Tutorial
48
MMM Example – Data Access(3)
c) employ loop blocking and interchange
loop i {loop k {loop j {
c[i][j] = c[i][j] + a[i][k] * b[k][j];
}}}→
loop k step s {loop j step s {loop i {
for (kk = k; kk < k+s; kk++) {for (jj = j; jj < j+s; jj++) {
c[i][jj] = c[i][jj] + a[i][kk] * b[kk][jj];
}}
}}}
Why?
PerfExpert Tutorial
49
MMM Example – Data Access(4)
Blocked loop code (blocking factor s = 70)
for (k = 0; k < n; k += s) {
for (j = 0; j < n; j += s) {
for (i = 0; i < n; i++) {
for (kk = k; kk < k + s; kk++) {
for (jj = j; jj < j + s; jj++) {
c[i][jj] += a[i][kk] * b[kk][jj];
}
}
}
}
}
PerfExpert Tutorial
50
MMM Example – Data Access(5)
total runtime in ./hpctoolkit-mmm3-database-1495284/experiment.xml is 0.28 s
Suggestions on how to alleviate performance bottlenecks are available at:
http://www.tacc.utexas.edu/perfexpert/
...
loop at line 28 in main (98.8% of the total runtime)
----------------------------------------------------------------------------performance assessment
LCPI good......okay......fair......poor......bad....
- overall
0.6 >>>>>>
upper bound by category
- data accesses
2.1 >>>>>>>>>>>>>>>>>>>>>
- instruction accesses
0.6 >>>>>>
- data TLB
0.0 >
- instruction TLB
0.0 >
- branch instructions
0.2 >>
- floating-point instr
2.5 >>>>>>>>>>>>>>>>>>>>>>>>>
PerfExpert Tutorial
51
Heat Transfer Code Example
Code Structure
Execution Example
PerfExpert Run
Optimization
PerfExpert Tutorial
52
Flow Chart
Heat Conduction Model
Input
Define Array
Grid generation
Initialize
Update
Make Matrix
TDMA
BCT
BCT
No
iter > max_iter
Yes
Plot
No
time > max_iter_time
PerfExpert Tutorial
Yes
53
Heat Transfer Code
Cartesian coordinate index extension for 3D mesh
Extra works in main.f90 and exchange.f90
Virtual (cartesian) topology implementation
MPI data type setup in main.f90
2-way communication routines in exchange.f90
Other processes are the same as the 2-D code
PerfExpert Tutorial
54
Output of Original Code
## CASE #1 K-I-J Loop ##
========================================================
Number of MPI processes is : 32
========================================================
# of iteration / comp. time / comm. time / total runtime
========================================================
100 10.08139 3.31155 10.39492
200 20.14803 6.61543 20.77528
300 30.20992 9.91811 31.15046
400 40.27607 13.22588 41.52979
500 50.33679 16.52017 51.90433
600 60.39856 19.82516 62.28027
700 70.46096 23.12081 72.65577
800 80.52318 26.42411 83.03120
900 90.58822 29.73077 93.40942
1000 100.64964 33.02405 103.78358
PerfExpert Tutorial
55
PerfExpert Output on Original Code
total runtime in heat3d-mpi-kij.exe is 100.95 seconds
Suggestions on how to alleviate performance bottlenecks are available at:
http://www.tacc.utexas.edu/perfexpert/
tdma (65.6% of the total runtime)
----------------------------------------------------------------------performance assessment
LCPI good......okay......fair......poor......bad....
- overall
6.5 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
upper bound by category
- data accesses
14.9 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
- instruction accesses
0.4 >>>>
- data TLB
0.0 >
- instruction TLB
0.0 >
- branch instructions
0.1 >
- floating-point instr
1.2 >>>>>>>>>>>>
loop at line 54 in tdma (52.1% of the total runtime)
PerfExpert Tutorial
56
Loop Structure in Original Code
do K = SZ1, EZ1
do I = SX1, EX1
do J = SY1, EY1
D(I,J,K)=Sterm(I,J,K)+A_E(I,J,K)*PHI_OLD(I+1,J,K) &
+A_W(I,J,K)*PHI_OLD(I-1,J,K) &
+A_T(I,J,K)*PHI_OLD(I,J,K+1) &
+A_B(I,J,K)*PHI_OLD(I,J,K-1)
TDMA_P(I,J,K)=A_N(I,J,K)/(A_P(I,J,K)A_S(I,J,K)*TDMA_P(I,J-1,K))
TDMA_Q(I,J,K)=(D(I,J,K)+A_S(I,J,K)*TDMA_Q(I,J1,K))/(A_P(I,J,K)A_S(I,J,K)*TDMA_P(I,J-1,K))
enddo
enddo
PerfExpert Tutorial
57
What is Wrong with Loop Structure?
For multi-dimensioned arrays, access is faster if
you iterate on the array subscript offering the
smallest stride or step size. In Fortran programs,
this is the leftmost subscript due to its columnmajor memory access pattern.
PerfExpert suggestion
Loop blocking or loop interchange
Interchange is simpler – let’s try it first
PerfExpert Tutorial
58
Code after Loop Index Interchange
do K = SZ1, EZ1
do J = SY1, EY1
do I = SX1, EX1
D(I,J,K)=Sterm(I,J,K)+A_E(I,J,K)*PHI_OLD(I+1,J,K) &
+A_W(I,J,K)*PHI_OLD(I-1,J,K) &
+A_T(I,J,K)*PHI_OLD(I,J,K+1) &
+A_B(I,J,K)*PHI_OLD(I,J,K-1)
TDMA_P(I,J,K)=A_N(I,J,K)/(A_P(I,J,K)-A_S(I,J,K)*
TDMA_P(I,J-1,K))
TDMA_Q(I,J,K)=(D(I,J,K)+A_S(I,J,K)*TDMA_Q(I,J-1,K))/
(A_P(I,J,K)- A_S(I,J,K)*TDMA_P(I,J-1,K))
enddo
enddo
PerfExpert Tutorial
59
PerfExpert on Interchanged Loop
total runtime in heat3d-mpi-kji.exe is 76.93 seconds
Suggestions on how to alleviate performance bottlenecks are available at:
http://www.tacc.utexas.edu/perfexpert/
tdma (73.4% of the total runtime)
----------------------------------------------------------------------------performance assessment
LCPI good......okay......fair......poor......bad....
- overall
6.2 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
upper bound by category
- data accesses
5.4 >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>+
- instruction accesses
0.4 >>>>
- data TLB
0.0 >
- instruction TLB
0.0 >
- branch instructions
0.1 >
- floating-point instr
1.6 >>>>>>>>>>>>>>>>
loop at line 53 in tdma (54.7% of the total runtime)
PerfExpert Tutorial
60
Summary and Future of Optimization
Optimization mostly still requires human pattern
recognition skills
Experience is the best guide to pattern recognition
PerfExpert will add optimization case studies as
they are generated
We hope you will send us case studies based on
your own codes and applications of PerfExpert
We hope to automate certain common
optimizations
PerfExpert Tutorial
61
Overview
Introduction
Why yet another performance tool?
What PerfExpert does and how it works
Demonstration
Quick-start example and hands-on demo
Optimization
Optimization process
Code tuning examples
Foundations
Porting and installation
PerfExpert Tutorial
62
Porting and Installing Outline
System requirements
Software stack
PerfExpert configuration
Establish categories, define LCPI formulae, select
performance counters, measure system parameters
PerfExpert installation
Example: AMD Barcelona configuration (Ranger)
Demo: porting to Intel Nehalem (Longhorn)
PerfExpert Tutorial
63
System Requirements for PerfExpert
Operating system
Linux
Software
Next slide
Compilers
C/C++ (for building tools and libraries)
Architecture
x86_64 (currently AMD Barcelona and Intel Nehalem)
PerfExpert Tutorial
64
Software Stack for PerfExpert
PerfCtr patch
http://user.it.uu.se/~mikpe/linux/perfctr/
PAPI library
http://icl.cs.utk.edu/papi/
HPCToolkit
http://hpctoolkit.org/
Java and Perl
http://www.java.com/
http://www.perl.org/
PerfExpert Tutorial
65
PerfExpert Configuration
1. Establish desired categories
Each bar in the output corresponds to a category
2. Define LCPI formula for each category
3. Select hardware performance counters
4. Determine system parameters
Measure the necessary system properties
The categories, formulae, counters, and parameters
that follow are appropriate for Ranger
PerfExpert Tutorial
66
1. Establish Desired Categories
Each bar in the output corresponds to a category
Currently, the following categories are used
Overall LCPI
Upper LCPI bounds
Data accesses
Instruction accesses
Data TLB
Instruction TLB
Branch instructions
Floating-point instructions
PerfExpert Tutorial
67
2. Define LCPI Formulae
Overall LCPI
TOT_CYC / TOT_INS
Upper LCPI bounds by category
(L1_DCA * L1_dlat + L2_DCA * L2_lat + L2_DCM * Mem_lat) / TOT_INS
(L1_ICA * L1_ilat + L2_ICA * L2_lat + L2_ICM * Mem_lat) / TOT_INS
(TLB_DM * TLB_lat) / TOT_INS
(TLB_IM * TLB_lat) / TOT_INS
(BR_INS * BR_lat + BR_MSP * BR_miss_lat) / TOT_INS
((FML_INS + FAD_INS) * FP_lat + FDV_INS * FP_slow_lat) / TOT_INS
PerfExpert Tutorial
68
3. Select Hardware Counters
Determine available performance counters
papi_avail and papi_native_avail
Select counters and test with microbenchmarks
CPU cycles (TOT_CYC)
Data TLB misses (TLB_DM)
Instructions committed (TOT_INS) Instruction TLB misses (TLB_IM)
L1 data cache accesses (L1_DCA) Branch instructions (BR_INS)
L2 cache data accesses (L2_DCA) Branch mispredictions (BR_MSP)
L2 cache data misses (L2_DCM)
Floating-point add/sub (FAD_INS)
L1 instr cache accesses (L1_ICA)
Floating-point mul (FML_INS)
L2 cache instr accesses (L2_ICA)
Floating-point div/sqrt (FDV_INS)
L2 cache instr misses (L2_ICM)
PerfExpert Tutorial
69
4. Determine System Parameters
For parameters with little or no variability, look
up or measure constant/worst-case value
CPU frequency
FP add/sub/mul latency (FP_lat)
L1 data cache latency (L1_dlat) FP div/sqrt latency (FP_slow_lat)
L1 instr cache latency (L1_ilat)
Branch instruction latency (BR_lat)
L2 cache latency (L2_lat)
Branch penalty (BR_miss_lat)
For parameters with large range, use
conservative value (may have to be tuned)
TLB miss latency (TLB_lat)
Memory access latency (Mem_lat)
Good CPI threshold
PerfExpert Tutorial
70
PerfExpert on Ranger
Demo: Ranger configuration
Demo: PerfExpert installation
PerfExpert Tutorial
71
Porting PerfExpert to Intel Nehalem
Software stack is already installed
Demonstration
PerfExpert configuration steps
Same categories as on Ranger
Performance counter selection
System parameters
LCPI formulae
PerfExpert file generation
PerfExpert.sge , PerfExpert.perl, and README
PerfExpert Tutorial
72
Porting from Ranger to Longhorn
1. Establish desired categories
No change
2. Define LCPI formula for each category
Adjust data accesses, instr accesses, FP instructions
3. Select hardware performance counters
Remove L2_DCM, FAD_INS, FML_INS, FDV_INS
Add L2_TCM, FP_COMP_OPS_EXE, ARITH, L1I
4. Determine system parameters
Same parameters (minus one), mostly different values
PerfExpert Tutorial
73
Longhorn LCPI Formulae
Overall LCPI
TOT_CYC / TOT_INS
Upper LCPI bounds by category
(L1_DCA * L1_dlat + L2_DCA * L2_lat + (L2_TCM - L2_ICM)* Mem_lat) / TOT_INS
(L1_ICA * L1_ilat + L1I + L2_ICA * L2_lat + L2_ICM * Mem_lat) / TOT_INS
(TLB_DM * TLB_lat) / TOT_INS
(TLB_IM * TLB_lat) / TOT_INS
(BR_INS * BR_lat + BR_MSP * BR_miss_lat) / TOT_INS
(FP_COMP_OPS_EXE * FP_lat + ARITH) / TOT_INS [no FP_slow_lat]
PerfExpert Tutorial
74
PerfExpert on Longhorn
Demo: Ranger configuration
Demo: PerfExpert installation
Future work
PerfExpert Tutorial
75
Future Work (1)
More case studies
Applications with various bottlenecks to harden tool
Improve and expand capabilities
Finer-grained recommendations
Add data structure based analyses and optimizations
Non-performance-counter-based measurements
Grow optimization and example database
Automatic implementation of solutions to common
core, chip and node-level performance bottlenecks
PerfExpert Tutorial
76
Future Work (2)
Port and deploy PerfExpert on other systems
New generation AMD and Intel chips
PowerPC chips
We solicit collaborations for other chip/node
architectures
Long term
Heterogeneous architectures
Communication and I/O bottleneck analyses and
optimizations
PerfExpert Tutorial
77
Goals for Tutorial
Operational effectiveness with PerfExpert
Be comfortable using PerfExpert
See how easy it is to use
Understanding of how PerfExpert works
Appreciate the sophistication of the analysis engine
Solicit collaborations
For applications of PerfExpert
For installation of PerfExpert at participant-local sites
Seek feedback on all aspects
PerfExpert Tutorial
78