phd- seminar . ppt - Uppsala University

Download Report

Transcript phd- seminar . ppt - Uppsala University

Dissertation Seminar, 18/11 – 2005
Auditorium Minus, Museum Gustavianum
Software Techniques for
Distributed Shared Memory
Zoran Radovic
[email protected]
[email protected]
Dissertation Seminar
Nov 18, 2005
Outline
 NUCA Locks
 DSZOOM – Software-based Shared Memory
 TMA – Trap-based Memory Architecture
[email protected]
Dissertation Seminar
Nov 18, 2005
Vasaloppet
“Contention Problem in Sweden”
Traditional cross-country ski race
90 km …
[email protected]
Dissertation Seminar
85.6533
km to
go…
CS
Nov 18, 2005
Critical Section (CS) Cost
Spin Locks under Contention
Spin locks
Spin locks
with backoff
IF (more contention) 
THEN less efficient CS …
“The more important the slower it runs…”
Amount of Contention
[email protected]
Dissertation Seminar
Nov 18, 2005
Queue-based Locks
CS Cost
Spin locks
Spin locks
with backoff
locks
IF (more contention)Queue-based

THEN constant CS cost …
Amount of Contention
[email protected]
Dissertation Seminar
Nov 18, 2005
CS Cost
This Dissertation
Spin locks
Spin locks
with backoff
IF (more contention) 
THEN more efficient CS …
Queue-based locks
“The more important the faster
it runs…”
NUCA locks
Amount of Contention
[email protected]
Dissertation Seminar
Nov 18, 2005
NUCA Locks (Basic Idea)
1) Reduce Switch
traffic
- one CPU per node is testing…
Memory
$
$
P
P
…
2) ImproveMemory
lock handover
3) More efficient CS
$
P
Lock/Unlock
[email protected]
- local traffic is cheaper
$
$
P
P
…
Memory
$
$
$
P
P
P
Test
Test
Test
Test
Lock/Unlock
Dissertation Seminar
…
$
P
Test
Test
Test
Test
Test
Test
Test
Nov 18, 2005
The HBO Lock (the simplest HBO)
 What do we need?
 node_id
Creates
 Compare&swap (CAS)
atomic operation
Communication
CAS(Lock_address, FREE, node_id)
 lock-acquire:
Affinity
 If the lock-value is in the state FREE:
• The node_id is CAS-ed into the lock location
 Else: 2 cases
• The lock is “local”  Spin with small backoff
• The lock is “remote”  Spin with large backoff
 Simple but fairly effective…
[email protected]
Dissertation Seminar
Nov 18, 2005
Performance Results
Realistic microbenchmark, 2-node WildFire, 28 CPUs
14
12
60
14
WF
50
Spin
MCS
HBO
10
9
Node Handoffs [%]
Iteration Time [seconds]
11
40
Fairness?
30
8
7
6
5
20
10
4
0
3
0
500
1000
1500
critical_work
[email protected]
2000
0
500
1000
1500
2000
critical_work
Dissertation Seminar
Nov 18, 2005
Fairness Study
Number of Finished Processors
Realistic microbenchmark, 2-node WildFire, 28 CPUs
t
28
26
24
22
20
18
16
14
12
10
8
6
4
2
0
Spin
MCS
HBO
0
[email protected]
5
Time [seconds]
Dissertation Seminar
10
15
Nov 18, 2005
Application Performance
28-processor runs
MCS
Spin EXP
Spin
Normalized Speedup
2.5
HBO
≈ 4x
2
1.5
1
0.5
[email protected]
Dissertation Seminar
ge
er
a
at
er
W
Av
-N
sq
nd
lre
Vo
ce
R
ay
t ra
io
si
ty
R
ad
FM
M
le
sk
y
C
ho
Ba
rn
es
0
Nov 18, 2005
Total Traffic: Raytrace
Local Transactions
Global Transactions
1.4
1.2
1
0.8
0.6
0.4
0.2
0
Spin
[email protected]
Spin EXP
Dissertation Seminar
MCS
HBO
Nov 18, 2005
HBO Locks inside Linux Kernel
 Patch provided by Silicon Graphics, Inc.
 Linux-IA64 kernel implementation, May 2005
 Page-fault handler runs 3x faster
 60 processors
[email protected]
Dissertation Seminar
Nov 18, 2005
Outline
 NUCA Locks
 DSZOOM – Software-based Shared Memory
 TMA – Trap-based Memory Architecture
[email protected]
Dissertation Seminar
Nov 18, 2005
The DSZOOM Proposal
[email protected]
Dissertation Seminar
Nov 18, 2005
The DSZOOM Proposal
 Run entire protocol in requesting-processor
 No protocol agent communication!
 Assumes user-level remote memory access
 put, get, and atomics [  InfiniBand ]
 Fine-grain memory protocols (e.g., 64 bytes)
 Hardware-like memory models
[Shasta, Blizzard, Sirocco]
[email protected]
Dissertation Seminar
Nov 18, 2005
“Squeezing” Protocols into Binaries…
DSZOOM
Program
Original
Program
...
cmp
bne
nop
...
cmp
bne
nop
%g0, %l5
0x24431
ld
[%o1 + 64], %o0
ldd
clr
...
[%o0 + 16], %f4
%l5
%g0, %l5
0x24431
ld
mov
and
cmp
bne
nop
[%o1 + 64], %o0
255, %g6
%g6, %o0, %g6
%g6, 170
0x24450
ldd
clr
...
[%o0 + 16], %f4
%l5
Fast-path
Protocol
Code
Slow-path
Protocol
Code
(C-code)
Binary/Assembler level instrumentation
[email protected]
Dissertation Seminar
Nov 18, 2005
Write Permission Caching
 Problem: store instrumentation relies on locking
 More complex instrumentation
 Solution: write permission cache (WPC)
 Small and fast software-managed cache
 Keeps write permissions
 The WPC idea:
 Exploit store locality
 Dynamically reduce the number of memory references
in store checking code
[email protected]
Dissertation Seminar
Nov 18, 2005
Other “Features”
 Two kinds of protocols
 Invalidate
 Update
 Many optimizations







Instrumentation scheduling
Instrumentation batching
WPC-based write batching
WPC-based dirty-data filtering
Private-data filtering
# of WPC entries
Coherence unit size
[email protected]
Dissertation Seminar
(update and invalidate)
(invalidate)
(update)
(update)
(update)
(update and invalidate)
(update and invalidate)
Nov 18, 2005
Coherence Flags and Profiling
 Coherence flags
 Similar to optimization flags of compilers
 Possible scenario:
gcc -dszoom-cl 128 -dszoom-inv –O3 my_app.c
 Execution profiling
 Similar to profile feedback of compilers
 Helps finding appropriate coherence flag settings
 Low overhead implementation in DSZOOM
• Less than 30 percent overhead
 Works for both small and large input sets
[email protected]
Dissertation Seminar
Nov 18, 2005
DSZOOM Results
2-node WildFire, 16 CPUs
[email protected]
Dissertation Seminar
BEST
1.11x
ra
di
os
ity
ra
yt
ra
ce
w
at
er
-n
sq
w
at
er
-s
p
an
-n
c
oc
e
an
-c
oc
e
fm
m
ba
rn
es
ra
di
x
lu
-n
c
lu
-c
1.45x
PROFILED
ra
ge
inv-dwpc-64
av
e
inv-64
3.0
2.8
2.6
2.4
2.2
2.0
1.8
1.6
1.4
1.2
1.0
0.8
0.6
0.4
0.2
0.0
fft
Normalized Execution Time
HW-DSM
Nov 18, 2005
Outline
 NUCA Locks
 DSZOOM – Software-based Shared Memory
 TMA – Trap-based Memory Architecture
[email protected]
Dissertation Seminar
Nov 18, 2005
Instrumentation Drawbacks
DSZOOM
Program
Original
Program
...
cmp
bne
nop
...
cmp
bne
nop
%g0, %l5
0x24431
ld
[%o1 + 64], %o0
ldd
clr
...
[%o0 + 16], %f4
%l5
•
•
ld
mov
and
cmp
bne
nop
[%o1 + 64], %o0
255, %g6
%g6, %o0, %g6
%g6, 170
0x24450
ldd
clr
...
[%o0 + 16], %f4
%l5
Binary transparency?
Run-time execution overhead
[email protected]
%g0, %l5
0x24431
Dissertation Seminar
Fast-path
Protocol
Code
Slow-path
Protocol
Code
(C-code)
Nov 18, 2005
Trap-Based Memory Architectures
 Basic idea
 Detect fine-grained coherence violations in hardware
 Trigger a coherence trap when one occur
 Maintain coherence by software protocols
 No memory system modifications
 Minimal processor modifications
 Binary Transparency
 No need to instrument binaries/applications
[email protected]
Dissertation Seminar
Nov 18, 2005
TMA Lite
Proof-of-concept Implementation
 Load permission check
 Hardware implementation of software check
• Predefined “magic-value” convention
 Store permission check
 Hardware WPC
 Can be seen as a very small cache
 Operates on virtual addresses
 Accessed in parallel with the data TLB
[email protected]
Dissertation Seminar
Nov 18, 2005
TMA Lite Performance
[TMA: simulation study, 4 nodes | DSZOOM: 2-node WildFire]
HW-DSM
DSZOOM
DWPC
PROFILED
BEST
TMA
Normalized Execution Time
2.5
1.75x
2
1.01x
1.5
1
0.5
[email protected]
Dissertation Seminar
av
er
ag
e
at
er
-s
p
w
at
er
-n
sq
w
ra
di
x
lu
-n
c
lu
-c
fft
0
Nov 18, 2005
Topics not Presented
 RH lock algorithm
 Controlled (un)fairness
 HBO_GT and HBO_GT_SD algorithms
 Global throttling and starvation detection
 DSZOOM implementation details
 Instrumentation challenges; scheduling, batching, etc.
 Bandwidth filtering techniques; dirty- & private-data
 Innovative TMA simulation tricks
 Low-level “good days” hacks
 Reusing Simics checkpoints
[email protected]
Dissertation Seminar
Nov 18, 2005
Dissertation Seminar, 18/11 – 2005
Auditorium Minus, Museum Gustavianum
Software Techniques for
Distributed Shared Memory
Zoran Radovic
[email protected]
[email protected]
Dissertation Seminar
Nov 18, 2005