Transcript ppt 1

Systems I
Pipelining IV
Topics


Implementing pipeline control
Pipelining and performance analysis
Implementing Pipeline Control
W
icode
valE
valM
dstE dstM
valE
valA
dstE dstM
M_icode
M
icode
Bch
e_Bch
CC
CC
E_dstM
E_icode
Pipe
control
logic
E_bubble
E
icode ifun
valC
valA
valB
dstE dstM srcA srcB
d_srcB
d_srcA
srcB
D_icode
srcA
D_bubble


D_stall
D
F_stall
F
icode ifun
rA
rB
valC
valP
predPC
Combinational logic generates pipeline control signals
Action occurs at start of following cycle
2
Initial Version of Pipeline Control
bool F_stall =
# Conditions for a load/use hazard
E_icode in { IMRMOVL, IPOPL } && E_dstM in { d_srcA, d_srcB } ||
# Stalling at fetch while ret passes through pipeline
IRET in { D_icode, E_icode, M_icode };
bool D_stall =
# Conditions for a load/use hazard
E_icode in { IMRMOVL, IPOPL } && E_dstM in { d_srcA, d_srcB };
bool D_bubble =
# Mispredicted branch
(E_icode == IJXX && !e_Bch) ||
# Bubble for ret
IRET in { D_icode, E_icode, M_icode };
bool E_bubble =
# Mispredicted branch
(E_icode == IJXX && !e_Bch) ||
# Load/use hazard
E_icode in { IMRMOVL, IPOPL } && E_dstM in { d_srcA, d_srcB};
3
Control Combinations
Load/use
M
E
D
Load
Use
ret 1
Mispredict
M
E
D
JXX
M
E
D
ret
ret 2
M
E
D
ret
bubble
ret 3
M
E
D
ret
bubble
bubble
Combination A
Combination B

Special cases that can arise on same clock cycle
Combination A


Not-taken branch
ret instruction at branch target
Combination B


Instruction that reads from memory to %esp
Followed by ret instruction
4
Control Combination A
ret 1
Mispredict
M
M
JXX
E
D
E
D
ret
Combination A
Condition
F
D
E
M
W
stall
bubble
normal
normal
normal
Mispredicted Branch normal
bubble
bubble
normal
normal
Combination
bubble
bubble
normal
normal
Processing ret



stall
Should handle as mispredicted branch
Stalls F pipeline register
But PC selection logic will be using M_valM anyhow
5
Stall in F
Your book provides two inconsistent meanings for
“stall in F”
Instruction remains in F and injects a bubble into D
Instruction squashed into D, same PC fetched
Figure 4.61
F
D
E
M
W
F
D
E
M
W
Use the one that keeps 1 instr per pipeline stage
6
JXX + ret works great!
1
2
3
4
5
6
7
8
9
0x000:
xorl %eax,%eax
F
D
E
M
W
0x002:
jne target # Not taken
F
D
E
M
W
F
D
E
M
W
bubble
D
E
M
W
0x007:
irmovl $1,%eax # Fall through
F
D
E
M
W
0x00d:
nop
F
D
E
M
0x011: t: ret
# Target
bubble
0x012:
nop
# Target + 1
10
F
W
7
Control Combination B
ret 1
Load/use
M
M
E
Load
E
D
Use
D
ret
Combination B
Condition
F
D
E
M
W
Processing ret
stall
bubble
normal
normal
normal
Load/Use Hazard
stall
stall
bubble
normal
normal
Combination
stall
bubble + bubble
stall
normal
normal


Would attempt to bubble and stall pipeline register D
Signaled by processor as pipeline error
8
Handling Control Combination B
ret 1
Load/use
M
M
E
Load
E
D
Use
D
ret
Combination B
Condition
F
D
E
M
W
Processing ret
stall
bubble
normal
normal
normal
Load/Use Hazard
stall
stall
bubble
normal
normal
Combination
stall
stall
bubble
normal
normal


Load/use hazard should get priority
ret instruction should be held in decode stage for additional
cycle
9
Corrected Pipeline Control Logic
bool D_bubble =
# Mispredicted branch
(E_icode == IJXX && !e_Bch) ||
# Stalling at fetch while ret passes
IRET in { D_icode, E_icode, M_icode
# but not condition for a load/use
&& !(E_icode in { IMRMOVL, IPOPL }
&& E_dstM in { d_srcA, d_srcB
Condition
through pipeline
}
hazard
});
F
D
E
M
W
Processing ret
stall
bubble
normal
normal
normal
Load/Use Hazard
stall
stall
bubble
normal
normal
Combination
stall
stall
bubble
normal
normal


Load/use hazard should get priority
ret instruction should be held in decode stage for additional
cycle
10
Load/use hazard with ret
mrmovl
ret
F D
F
mrmovl F D E
ret
F D
addl
F
mrmovl F D E M
bubble
E
ret
F D D
addl
F F
mrmovl F D E M W
bubble
E M
ret
F D D E
addl
F F
bubble
D
addl
F
11
Pipeline Summary
Data Hazards

Most handled by forwarding
 No performance penalty

Load/use hazard requires one cycle stall
Control Hazards

Cancel instructions when detect mispredicted branch
 Two clock cycles wasted

Stall fetch stage while ret passes through pipeline
 Three clock cycles wasted
Control Combinations


Must analyze carefully
First version had subtle bug
 Only arises with unusual instruction combination
12
Performance Analysis with Pipelining
CPU time 
Seconds Instructio ns
Cycles
Seconds



Program
Program
Instructio n Cycle
Ideal pipelined machine: CPI = 1


One instruction completed per cycle
But much faster cycle time than unpipelined machine
However - hazards are working against the ideal


Hazards resolved using forwarding are fine
Stalling degrades performance and instruction comletion
rate is interrupted
CPI is measure of “architectural efficiency” of design
13
CPI for PIPE
CPI  1.0


Fetch instruction each clock cycle
Effectively process new instruction almost every cycle
 Although each individual instruction has latency of 5 cycles
CPI > 1.0

Sometimes must stall or cancel branches
Computing CPI




C clock cycles
I instructions executed to completion
B bubbles injected (C = I + B)
CPI = C/I = (I+B)/I = 1.0 + B/I
Factor B/I represents average penalty due to bubbles
14
Computing CPI
CPI

Function of useful instruction and bubbles
CPI 

Ci  Cb
C
 1.0  b
Ci
Ci
Cb/Ci represents the pipeline penalty due to stalls
Can reformulate
to account for




load penalties (lp)
branch misprediction penalties (mp)
return penalties (rp)
CPI 1.0  lp  mp rp

15
Computing CPI - II
So how do we determine the penalties?




Depends on how often each situation occurs on average
How often does a load occur and how often does that load
cause a stall?
How often does a branch occur and how often is it
mispredicted
How often does a return occur?
We can measure these


simulator
hardware performance counters
We can estimate through historical averages

Then use to make early design tradeoffs for architecture
16
Computing CPI - III
Cause
Name
Instruction Condition
Frequency Frequency
Stalls
Product
Load/Use
lp
0.30
0.3
1
0.09
Mispredict
mp
0.20
0.4
2
0.16
Return
rp
0.02
1.0
3
0.06
Total penalty
0.31
CPI = 1 + 0.31 = 1.31 == 31% worse than ideal
This gets worse when:


Account for non-ideal memory access latency
Deeper pipelines (where stalls per hazard increase)
17
CPI for PIPE (Cont.)

B/I = LP + MP + RP
LP: Penalty due to load/use hazard stalling
 Fraction of instructions that are loads
 Fraction of load instructions requiring stall
 Number of bubbles injected each time
Typical Values
0.25
0.20
1
 LP = 0.25 * 0.20 * 1 = 0.05

MP: Penalty due to mispredicted branches
 Fraction of instructions that are cond. jumps
 Fraction of cond. jumps mispredicted
 Number of bubbles injected each time
0.20
0.40
2
 MP = 0.20 * 0.40 * 2 = 0.16

RP: Penalty due to ret instructions
 Fraction of instructions that are returns
 Number of bubbles injected each time
0.02
3
 RP = 0.02 * 3 = 0.06

Net effect of penalties 0.05 + 0.16 + 0.06 = 0.27
 CPI = 1.27
(Not bad!)
18
Summary
Today


Pipeline control logic
Effect on CPI and performance
Next Time


Further mitigation of branch mispredictions
State machine design
19