Transcript powerpoint

Emulating Unimplemented
Instructions in an SMT
Suan Yong & Brian Forney
CS/ECE 752
Spring 2000
Motivation
• Simultaneous Multithreaded processors are
promising and likely to be embraced by
industry
– Exploit thread level parallelism
– Compaq’s Alpha 21464 has planned SMT
support
The question
• What if SMT support was needed but cost,
complexity, or power consumption are an
issue?
• One solution is to remove functional units
=> emulation
• Can anything be done to speed up emulation
of instructions?
Related work
• “The Use of Multithreading for Exception
Handling,” Zilles et al, Micro-32,
November 1999
• “Simultaneous Subordinate Microthreading
(SSMT),” Chappell et al, 26th Annual
ISCA, May 1999
Exception
Handling
standard
pipeline:
3
4
5
6
6
7
7
A
B
C
C
6
7
8
3
4
5
6
7
SMT:
A
B
C
C
8
9
10
BRANCH PREDICTOR
FETCH
I$
3 1 2 1
DECODE
3 1 2 1
PC
R1
R2
:
Rn
R1
R2
:
Rn
I-MUL
3 2 1
I-ALU
FP
R/W
D$
PC
Simultaneous
Multithreading
(SMT)
1
Emulating SMT approach
BRANCH PREDICTOR
FETCH
I$
DECODE
PC
R1
R2
:
Rn
R1
R2
:
Rn
T-RET
T-STRT
I-MUL
I-ALU
FP
R/W
D$
PC
BRANCH PREDICTOR
FETCH
I$
B A
DECODE
7 6 5 4
&
R1
R2
:
Rn
R1
R2
:
Rn
src1
T-RET
5
T-STRT
[7] [6] [5] [4] [3] [2] [1]
PC
I-MUL
7 6 5 4 3
I-ALU
FP
R/W
D$
A
PC
src2
[3]
BRANCH PREDICTOR
FETCH
I$
Z PC
Z Z Z
ZR1 Z Z
Z R2
Z Z Z
Z: Z Z
Z Rn
Z Z Z
C B A
DECODE
A 7 6 5 4
R1
R2
:
Rn
T-RET
5
T-STRT
[7] [6] [5] [4] [3] [2] [1]
I-MUL
7 6 5 4 3
I-ALU
FP
R/W
D$
PC
src1
src2
[3]
BRANCH PREDICTOR
FETCH
I$
Z PC
Z Z Z
ZR1 Z Z
Z R2
Z Z Z
Z: Z Z
Z Rn
Z Z Z
DECODE
6 C
R1
R2
:
Rn
C
T-RET
T-STRT
[7] [6] [5] [4] [3] [2] [1]
I-MUL
7 6 5 4 3
I-ALU
FP
R/W
D$
PC
C B A
src1
src2
[3]
BRANCH PREDICTOR
FETCH
I$
PC
R1
R2
:
Rn
DECODE
6 C
7 6 5 4 3
C
T-RET
T-STRT
I-MUL
I-ALU
FP
R/W
D$
Z PC
Z Z Z
ZR1 Z src1
Z
Z R2
Z Zsrc2Z
Z: Z [3]
Z
Z Rn
Z Z Z
C B A
[7] [6] [5] [4] [3] [2] [1]
?
Methodology
• Modified Zilles’s sim-multi Compaq Alpha
SMT simulator
– added exception thread support
– added multiply thread
• Ran representative execution traces of
benchmarks from SPEC CPU2000 and
MediaBench
Simulator modes
Baseline - no emulation
“Squash” - squash pipeline when emulation triggered
mode - mimics exception handler performance
without SMT
“Pause” - stop fetching when emulation triggered
mode - squash if thread is stalled (e.g. icache miss)
“ooo2” - original thread continues fetching after
“ooo4”
emulator is started
- emulator always has higher fetch priority
Emulation Slowdown (percentage)
(scaled to proportion of multiply instructions in program)
160
140
squash
pause
120
ooo2
ooo4
100
80
60
40
20
vo
rte
x
tw
ol
f
sw
im
se
r
pa
r
es
a
m
p
gz
i
ep
ic
dj
pe
g
cj
pe
g
cr
af
ty
0
Conclusions
• ESMT usually minimizes performance cost
of emulation
• “ooo” mode (non-pausing) works best
• “squash” is occasionally better, because of
resource contention
– do partial squashing?
• Some of the hardware is already needed,
and could be useful for other purposes