Transcript ppt
Using Interaction Cost (icost) for
Microarchitectural Bottleneck Analysis
Brian Fields1 Rastislav Bodik1
Mark Hill2
1UC-Berkeley, 2UW-Madison, 3Intel
Chris Newburn3
Outline
Interaction Cost
Bottleneck analysis complicated by parallelism
Parallelism causes interactions
• Qualitative: parallel and serial interactions
• Quantitative: interaction cost (icost)
Icost case study: designing a deep pipeline
Hardware profiler
Icost “shotgun” profiler
• Replace current performance counters
Bottleneck analysis is hard
Why?
-architectural parallelism complicates
performance understanding
• Two parallel cache misses
• A multiply and window stall
• A branch mispredict and full-store-buffer
stall occur in the same cycle that three loads
are waiting on the memory system and two
floating-point multiplies are executing
What we want from bottleneck analysis
Performance cost (or reward)
speedup when the bottleneck is removed
Q: What if two bottlenecks interact?
Our solution: measure interactions
Two parallel cache misses (Each 100 cycles)
miss #1 (100)
miss #2 (100)
Cost(miss #1) = 0
Cost(miss #2) = 0
Cost({miss #1, miss #2}) = 100
Aggregate cost > Sum of individual costs Parallel interaction
100
0+0
icost = aggregate cost – sum of individual costs
= 100 – 0 – 0 = 100
Interaction cost (icost)
icost = aggregate cost – sum of individual costs
1. Positive icost
parallel interaction
2. Zero icost ?
miss #1
miss #2
Interaction cost (icost)
icost = aggregate cost – sum of individual costs
1. Positive icost
parallel interaction
2. Zero icost
independent
3. Negative icost ?
miss #1
miss #1
miss #2
...
miss #2
Negative icost
Two serial cache misses (data dependent)
miss #1 (100) miss #2 (100)
ALU latency (110 cycles)
Cost(miss #1) = ?
Negative icost
Two serial cache misses (data dependent)
miss #1 (100) miss #2 (100)
ALU latency (110 cycles)
Cost(miss #1) = 90
Cost(miss #2) = 90
Cost({miss #1, miss #2}) = 90
icost = aggregate cost – sum of individual costs
= 90 – 90 – 90 = -90
Negative icost serial interaction
Interaction cost (icost)
icost = aggregate cost – sum of individual costs
1. Positive icost
miss #1
parallel interaction
2. Zero icost
independent
3. Negative icost
serial interaction
Fetch
BranchBW
mispredict
Load-Replay
LSQ
stall
Trap
miss #2
miss #1
..
.
miss #1
miss #2
miss #2
ALU latency
Why care about serial interactions?
miss #1 (100)
miss #2 (100)
ALU latency (110 cycles)
Reason #1 We are over-optimizing!
Prefetching miss #2 doesn’t help if miss #1 is
already prefetched (but the overhead still costs us)
Reason #2 We have a choice of what to optimize
Prefetching miss #2 has the same effect as miss #1
Icost Case Study: Deep pipelines
Deep pipelines cause long latency loops:
• level-one (DL1) cache access,
issue-wakeup, branch misprediction, …
But can often mitigate them indirectly
Assume 4-cycle DL1 access; how to mitigate?
Increase cache ports? Increase window size?
Increase fetch BW? Reduce cache misses?
Really, looking for serial interactions!
Icost Case Study: Deep pipelines
DL1 access
F
5
1
F
5
4
E
6
i1
0
F
1
C
i2
5
E
9
14
18
1
C
i3
0
F
5
E
4
C
12
5
2
E
C
i4
window edge
6
1
C
i5
F
5
E
E
7
0
12
F
1
7
0
C
i6
Icost Case Study: Deep pipelines
DL1 access
F
5
1
F
5
4
E
6
i1
0
F
1
C
i2
5
E
9
14
18
1
C
i3
0
F
5
E
4
C
12
5
2
E
C
i4
window edge
6
1
C
i5
F
5
E
E
7
0
12
F
1
7
0
C
i6
Icost Case Study: Deep pipelines
DL1 access
F
5
1
F
5
4
E
6
i1
0
F
1
C
i2
5
E
9
14
18
1
C
i3
0
F
5
E
4
C
12
5
2
E
C
i4
window edge
6
1
C
i5
F
5
E
E
7
0
12
F
1
7
0
C
i6
Icost Case Study: Deep pipelines
DL1 access
F
5
1
F
5
4
E
6
i1
0
F
1
C
i2
5
E
9
14
18
1
C
i3
0
F
5
E
4
C
12
5
2
E
C
i4
window edge
6
1
C
i5
F
5
E
E
7
0
12
F
1
7
0
C
i6
Icost Case Study: Deep pipelines
DL1 access
F
5
1
F
5
4
E
6
i1
0
F
1
C
i2
5
E
9
14
18
1
C
i3
0
F
5
E
4
C
12
5
2
E
C
i4
window edge
6
1
C
i5
F
5
E
E
7
0
12
F
1
7
0
C
i6
Icost Case Study: Deep pipelines
DL1 access
F
5
1
F
5
4
E
6
i1
0
F
1
C
i2
5
E
9
14
18
1
C
i3
0
F
5
E
4
C
12
5
2
E
C
i4
window edge
6
1
C
i5
F
5
E
E
7
0
12
F
1
7
0
C
i6
Icost Case Study: Deep pipelines
DL1 access
F
5
1
F
5
4
E
6
i1
0
F
1
C
i2
5
E
9
14
18
1
C
i3
0
F
5
E
4
C
12
5
2
E
C
i4
window edge
6
1
C
i5
F
5
E
E
7
0
12
F
1
7
0
C
i6
Icost Breakdown (6 wide, 64-entry window)
gcc
DL1
DL1+window
DL1+bw
DL1+bmisp
DL1+dmiss
DL1+alu
DL1+imiss
...
Total
gzip
vortex
Icost Breakdown (6 wide, 64-entry window)
gcc
DL1
DL1+window
DL1+bw
DL1+bmisp
DL1+dmiss
DL1+alu
DL1+imiss
...
Total
gzip
30.5 %
vortex
Icost Breakdown (6 wide, 64-entry window)
gcc
gzip
DL1
30.5 %
DL1+window
-15.3
DL1+bw
6.0
DL1+bmisp
-3.4
DL1+dmiss
-0.4
DL1+alu
-8.2
DL1+imiss
0.0
...
...
Total
100.0
vortex
Icost Breakdown (6 wide, 64-entry window)
gcc
gzip
vortex
DL1
18.3 %
30.5 %
25.8 %
DL1+window
-4.2
-15.3
-24.5
DL1+bw
10.0
6.0
15.5
DL1+bmisp
-7.0
-3.4
-0.3
DL1+dmiss
-1.4
-0.4
-1.4
DL1+alu
-1.6
-8.2
-4.7
DL1+imiss
0.1
0.0
0.4
...
...
...
...
Total
100.0
100.0
100.0
Vortex Breakdowns, enlarging the window
64
DL1
DL1+window
DL1+bw
DL1+bmisp
DL1+dmiss
DL1+alu
DL1+imiss
...
Total
128
256
Vortex Breakdowns, enlarging the window
64
128
256
DL1
25.8
8.9
3.9
DL1+window
-24.5
-7.7
-2.6
DL1+bw
15.5
16.7
13.2
DL1+bmisp
-0.3
-0.6
-0.8
DL1+dmiss
-1.4
-2.1
-2.8
DL1+alu
-4.7
-2.5
-0.4
DL1+imiss
0.4
0.5
0.3
...
...
...
...
Total
100.0
80.8
75.0
Outline
Interaction Cost
Bottleneck analysis complicated by parallelism
Parallelism causes interactions
• Qualitative: parallel and serial interactions
• Quantitative: interaction cost (icost)
Icost case study: designing a deep pipeline
• Exploiting serial interactions
Hardware profiler
Icost “shotgun” profiler
• Overcome the limitations of performance counters
Profiling goal
Goal:
•
Construct graph
many dynamic instructions
Constraint:
•
Can only sample sparsely
Genome sequencing
Profiling goal
Goal:
•
DNA strand
Construct graph
DNA
Constraint:
•
Can only sample sparsely
“Shotgun” genome sequencing
DNA
“Shotgun” genome sequencing
DNA
“Shotgun” genome sequencing
DNA
...
...
“Shotgun” genome sequencing
DNA
...
...
Find overlaps among samples
...
...
Mapping “shotgun” to our situation
many dynamic instructions
Icache miss
Dcache miss
Branch misp.
No event
Profiler hardware requirements
...
...
Profiler hardware requirements
...
Match!
...
Conclusion
Bottleneck analysis is complicated by parallelism
Parallelism is interpreted with interaction cost (icost)
• Three possibilities: independent, parallel, or serial
Applies to all instructions, resources, events
Enabled by the “shotgun” profiler:
Interaction cost overcomes limitations of counters
Icost Case Study: Deep pipelines
Decode,
rename
F
5
1
F
E
6
0
F
1
C
i2
Multiply +
pipe latency
5
E
9
14
18
1
C
i3
0
F
5
E
4
i1
12
5
4
C
Icache miss
DL1 access
5
2
E
C
i4
window edge
6
1
C
i5
F
5
E
E
7
0
12
F
1
7
0
C
i6
Profiler software requirements
Software puts the graph together
Detailed samples
(with matching PC)
Skeleton sample
Compare Icost and Sensitivity Study
Corollary to DL1 and ROB serial interaction:
As load latency increases, the benefit from enlarging
the ROB increases.
F
1
DL1 access
0
F
1
1
E
3
2
C
i1
0
C
i2
1
E
1
2
2
1
C
i3
1
F
1
4
E
F
0
1
2
E
C
i4
2
1
C
i5
F
1
E
E
3
0
1
F
1
3
0
C
i6
Compare Icost and Sensitivity Study
25
DL1 Latency
Speedup
20
10
5
4
3
2
1
15
10
5
0
64
128
192
ROB size
256
Compare Icost and Sensitivity Study
Sensitivity Study Advantages
•
More information
•
e.g., concave or convex curves
Interaction Cost Advantages
•
Easy (automatic) interpretation
•
•
Sign and magnitude have well defined meanings
Concise communication
•
DL1 and ROB interact serially