Transcript ppt

Systems I
Pipelining III
Topics



Hazard mitigation through pipeline forwarding
Hardware support for forwarding
Forwarding to mitigate control (branch)
hazards
How do we fix the Pipeline?
Pad the program with NOPs

Yuck!
Stall the pipeline

Data hazards
 Wait for producing instruction to complete
 Then proceed with consuming instruction

Control hazards
 Wait until new PC has been determined
 Then begin fetching

How is this better than putting NOPs into the program?
Forward data within the pipeline

Grab the result from somewhere in the pipe
 After it has been computed
 But before it has been written back

This gives an opportunity to avoid performance degradation due to
hazards!
2
Data Forwarding
Naïve Pipeline


Register isn’t written until completion of write-back stage
Source operands read from register file in decode stage
 Needs to be in register file at start of stage
Observation

Value generated in execute or memory stage
Trick


Pass value directly from generating instruction to decode
stage
Needs to be available at end of decode stage
3
Data Forwarding Example
# demo-h2.ys
1
2
3
4
5
0x000: irmovl $10,%edx
F
D
F
E
D
F
M
E
D
F
W
M
E
D
F
0x006: irmovl
$3,%eax
0x00c: nop
0x00d: nop
0x00e: addl %edx,%eax
0x010: halt



irmovl in writeback stage
Destination value in
W pipeline register
Forward as valB for
decode stage
6
7
8
9
10
W
M
E
D
F
W
M
E
D
W
M
E
W
M
W
Cycle 6
W
R[ %eax] f 3
W_dstE = %eax
W_valE = 3
•
•
•
D
srcA = %edx
srcB = %eax
valA f R[ %edx] = 10
valB f W_valE = 3
4
W_icode, W_valM
W_valE, W_valM, W_dstE, W_dstM
Bypass Paths
W_valE
W_valM
W
Decode Stage



m_valM
Forwarding logic selects
Memory
valA and valB
Normally from register
file
Forwarding: get valA or
valB from later pipeline
Execute
stage
Addr, Data
M_valE
M


Execute: valE
Memory: valE, valM
Write back: valE, valM
e_valE
Bch
CC
CC
ALU
ALU
E_valA, E_valB,
E_srcA, E_srcB
Forwarding Sources

Data
Data
memory
memory
M_icode,
M_Bch,
M_valA
E
valA, valB
Forward
d_srcA,
d_srcB
Decode
A
B
Register
Register M
file
file E
Write back
D
valP
5
Data Forwarding Example #2
# demo-h0.ys
1
2
3
4
5
6
0x000: irmovl $10,%edx
F
D
F
E
D
M
E
W
M
W
F
D
E
M
W
F
D
E
M
0x006: irmovl
$3,%eax
0x00c: addl %edx,%eax
0x00e: halt
Register %edx


Generated by ALU
during previous cycle
Forward from memory
as valA
Register %eax

Value just generated
by ALU

Forward from execute
as valB
7
8
W
Cycle 4
M
M_dstE = %edx
M_valE = 10
E
E_dstE = %eax
e_valE f 0 + 3 = 3
D
srcA = %edx
srcB = %eax
valA f M_valE = 10
valB f e_valE = 3
6
Implementing
Forwarding
W_valE
Write back
W_valM
W
icode
valE
valM
dstE dstM
data out
read
m_valM
Data
Data
memory
memory
Mem.
control
write
Memory

data in
Addr
M_Bch
M
icode
M_valA
M_valE
Bch
valE
valA
dstE dstM
e_Bch
e_valE
ALU
ALU
CC
CC
Execute
E
icode ifun
ALU
fun.
ALU
A
ALU
B
valC
valA
valB

Add additional feedback
paths from E, M, and W
pipeline registers into
decode stage
Create logic blocks to
select from multiple
sources for valA and valB
in decode stage
dstE dstM srcA srcB
d_srcA d_srcB
dstE dstM srcA srcB
Sel+Fwd
A
Decode
D
icode ifun
Fwd
B
A
W_valM
B
Register
Register M
file
file E
rA
rB
Instruction
Instruction
memory
valC
W_valE
valP
PC
PC
increment
Predict
PC
7
Implementing Forwarding
W_valE
W_valM
valE
valM
dstE dstM
data out
read
m_valM
Data
Data
memory
memory
l
write
data in
Addr
M_valA
M_valE
valE
valA
dstE dstM
e_valE
ALU
ALU
ALU
fun.
ALU
A
ALU
B
valC
valA
valB
dstE dstM srcA srcB
d_srcA d_srcB
dstE dstM srcA srcB
Sel+Fwd
A
Fwd
B
A
B
Register
Register M
file
file E
valC
valP
## What should be the A value?
int new_E_valA = [
# Use incremented PC
D_icode in { ICALL, IJXX } : D_valP;
# Forward valE from execute
d_srcA == E_dstE : e_valE;
# Forward valM from memory
d_srcA == M_dstM : m_valM;
# Forward valE from memory
d_srcA == M_dstE : M_valE;
# Forward valM from write back
d_srcA == W_dstM : W_valM;
# Forward valE from write back
d_srcA == W_dstE : W_valE;
# Use value read from register file
1 : d_rvalA;
];
W_valM
W_valE
8
Limitation of Forwarding
# demo-luh.ys
0x000:
0x006:
0x00c:
0x012:
0x018:
0x01e:
0x020:
1
2

4
irmovl $128,%edx
F
D
E M
irmovl $3,%ecx
F
D
E
rmmovl %ecx, 0(%edx)
F
D
irmovl $10,%ebx
F
mrmovl 0(%edx),%eax # Load %eax
addl %ebx,%eax # Use %eax
halt
Load-use dependency

3
Value needed by end of
decode stage in cycle 7
Value read from memory in
memory stage of cycle 8
5
6
W
M
E
D
F
W
M
E
D
F
7
8
9
10
11
W
M
E
D
F
W
M
E
D
W
M
E
W
M
W
Cycle 7
Cycle 8
M
M
M_dstE = %ebx
M_valE = 10
M_dstM = %eax
m_valM f M[128] = 3
•
•
•
D
valA f M_valE = 10
valB f R[%eax] = 0
Error
9
Avoiding Load/Use Hazard
# demo-luh.ys
1
2
3
4
5
irmovl $128,%edx
F
irmovl $3,%ecx
rmmovl %ecx, 0(%edx)
irmovl $10,%ebx
D
E
M
W
F
D
F
E
D
F
0x018: mrmovl 0(%edx),%eax # Load %eax
bubble
0x01e: addl %ebx,%eax # Use %eax
0x020: halt
M
E
D
F
0x000:
0x006:
0x00c:
0x012:


Stall using instruction for
one cycle
Can then pick up loaded
value by forwarding from
memory stage
6
7
8
9
W
M
E
D
W
M
E
W
M
W
D
F
E
D
F
M
E
D
F
10
11
W
M
E
W
M
12
W
Cycle 8
W
W_dstE = %ebx
W_valE = 10
M
M_dstM = %eax
m_valM f M[128] = 3
•
•
•
D
valA f W_valE = 10
valB f m_valM = 3
10
Data
Data
memory
memory
Mem.
control
write
Memory
Detecting Load/Use Hazard
data in
Addr
M_Bch
M
icode
M_valA
M_valE
Bch
valE
valA
dstE
dstM
e_Bch
e_valE
ALU
ALU
CC
CC
Execute
E
icode
ifun
ALU
fun.
ALU
A
ALU
B
valC
valA
valB
dstE
dstM
dstE
dstM
srcA
srcB
d_srcA d_srcB
Sel +Fwd
A
Decode
D
icode
rA
rB
valC
srcB
Fwd
B
A
ifun
srcA
W_valM
B
Register
RegisterM
file
file E
W_valE
valP
Predict
PC
Condition
Fetch
Instruction
Instruction
memory
memory
Trigger
PC
PC
increment
increment
f_PC
Load/Use Hazard
F
M_valA
E_icode in { IMRMOVL, IPOPL } &&
E_dstM in { d_srcA, d_srcB }
Select
PC
W_valM
predPC
11
Control for Load/Use Hazard
# demo-luh.ys
1
2
3
0x000:
0x006:
0x00c:
0x012:
0x018:
4
irmovl $128,%edx
F
D E M
irmovl $3,%ecx
F
D E
rmmovl %ecx, 0(%edx)
F
D
irmovl $10,%ebx
F
mrmovl 0(%edx),%eax # Load %eax
bubble
0x01e: addl %ebx,%eax # Use %eax
0x020: halt


5
6
7
W
M
E
D
F
W
M
E
D
W
M
E
F
D
F
8
9
10
11
W
M
E
D
F
W
M
E
D
W
M
E
W
M
12
W
Stall instructions in fetch
and decode stages
Inject bubble into execute
stage
Condition
Load/Use Hazard
F
D
E
M
W
stall
stall
bubble
normal
normal
12
Branch Misprediction Example
demo-j.ys
0x000:
xorl %eax,%eax
0x002:
jne t
0x007:
irmovl $1, %eax
0x00d:
nop
0x00e:
nop
0x00f:
nop
0x010:
halt
0x011: t: irmovl $3, %edx
0x017:
irmovl $4, %ecx
0x01d:
irmovl $5, %edx

# Not taken
# Fall through
# Target (Should not execute)
# Should not execute
# Should not execute
Should only execute first 7 instructions
13
Handling Misprediction
# demo-j.ys
1
2
3
4
5
6
0x000:
xorl %eax,%eax
F
0x002:
jne target # Not taken
D
F
E
D
F
M
E
D
W
M
W
E
M
W
D
F
E
D
F
0x011: t: irmovl $2,%edx # Target
bubble
0x017:
irmovl $3,%ebx # Target+1
bubble
0x007:
irmovl $1,%eax # Fall through
0x00d:
nop
7
8
9
M
E
W
M
W
D
E
M
10
F
W
Predict branch as taken

Fetch 2 instructions at target
Cancel when mispredicted



Detect branch not-taken in execute stage
On following cycle, replace instructions in execute and
decode by bubbles
No side effects have occurred yet
14
W_valM
W
icode
valE
valM
dstE
dstM
Detecting Mispredicted Branch
data out
read
Data
Data
memory
memory
Mem.
control
write
Memory
m_valM
data in
Addr
M_Bch
M
icode
M_valA
M_valE
Bch
valE
valA
dstE
dstM
e_Bch
e_valE
ALU
ALU
CC
CC
Execute
E
icode
ifun
ALU
fun.
ALU
A
ALU
B
valC
valA
valB
dstE
dstM
dstE
dstM
srcA
srcB
d_srcA d_srcB
Sel +Fwd
A
Condition
srcA
srcB
Fwd
B
Trigger
Decode
A
W_valM
B
Register
RegisterM
file
file E
Mispredicted Branch E_icode = IJXX & !e_Bch
D
Fetch
icode
ifun
rA
rB
Instruction
Instruction
memory
memory
valC
W_valE
valP
PC
PC
increment
increment
Predict
PC
f_PC
M_valA
Select
PC
W_valM
15
Control for Misprediction
# demo-j.ys
1
2
3
4
5
6
0x000:
xorl %eax,%eax
F
0x002:
jne target # Not taken
D
F
E
D
F
M
E
D
W
M
W
E
M
W
D
F
E
D
F
0x011: t: irmovl $2,%edx # Target
bubble
0x017:
irmovl $3,%ebx # Target+1
0x007:
irmovl $1,%eax # Fall through
0x00d:
nop
F
Mispredicted Branch normal
8
9
M
E
W
M
W
D
E
M
10
F
bubble
Condition
7
W
D
E
M
W
bubble
bubble
normal
normal
16
demo-retb.ys
Return Example
0x000:
0x006:
0x00b:
0x011:
0x020:
0x020:
0x026:
0x027:
0x02d:
0x033:
0x039:
0x100:
0x100:

irmovl Stack,%esp
call p
irmovl $5,%esi
halt
.pos 0x20
p: irmovl $-1,%edi
ret
irmovl $1,%eax
irmovl $2,%ecx
irmovl $3,%edx
irmovl $4,%ebx
.pos 0x100
Stack:
# Initialize stack pointer
# Procedure call
# Return point
# procedure
#
#
#
#
Should
Should
Should
Should
not
not
not
not
be
be
be
be
executed
executed
executed
executed
# Stack: Stack pointer
Previously executed three additional instructions
17
Correct Return Example
# demo-retb
0x026:
ret
F
bubble
D
E
M
W
F
D
E
M
W
F
D
E
M
W
F
D
E
M
W
F
D
E
M
bubble
bubble
0x00b:

irmovl $5,%esi # Return
As ret passes through
pipeline, stall at fetch stage
W
W
valM = 0x0b
 While in decode, execute, and
memory stage


Inject bubble into decode
stage
Release stall when reach
write-back stage
•
•
•
F
valC f 5
rB f %esi
18
Detecting Return
M_Bch
M
icode
M_valE
Bch
valE
valA
dstE dstM
e_Bch
e_valE
ALU
ALU
CC
CC
Execute
E
icode ifun
ALU
fun.
ALU
A
ALU
B
valC
valA
valB
dstE dstM srcA
srcB
d_srcA d_srcB
dstE dstM srcA
Sel+Fwd
A
Decode
D
icode ifun
Fwd
B
A
B
Register
Register M
file
file E
rA
rB
valC
srcB
W_valM
W_valE
valP
Condition
Trigger
Processing ret
IRET in { D_icode, E_icode, M_icode }
19
Control for Return
# demo-retb
0x026:
ret
F
bubble
D
E
M
W
F
D
E
M
W
F
D
E
M
W
F
D
E
M
W
F
D
E
M
bubble
bubble
0x00b:
irmovl $5,%esi # Return
Condition
Processing ret
W
F
D
E
M
W
stall
bubble
normal
normal
normal
20
Special Control Cases
Detection
Condition
Trigger
Processing ret
IRET in { D_icode, E_icode, M_icode }
Load/Use Hazard
E_icode in { IMRMOVL, IPOPL } &&
E_dstM in { d_srcA, d_srcB }
Mispredicted Branch E_icode = IJXX & !e_Bch
Action (on next cycle)
Condition
F
D
E
M
W
Processing ret
stall
bubble
normal
normal
normal
Load/Use Hazard
stall
stall
bubble
normal
normal
bubble
bubble
normal
normal
Mispredicted Branch normal
21
Summary
Today



Hazard mitigation through pipeline forwarding
Hardware support for forwarding
Forwarding to mitigate control (branch) hazards
Next Time


Implementing pipeline control
Pipelining and performance analysis
22