phd- seminar . ppt - Uppsala University
Download
Report
Transcript phd- seminar . ppt - Uppsala University
Dissertation Seminar, 18/11 – 2005
Auditorium Minus, Museum Gustavianum
Software Techniques for
Distributed Shared Memory
Zoran Radovic
[email protected]
[email protected]
Dissertation Seminar
Nov 18, 2005
Outline
NUCA Locks
DSZOOM – Software-based Shared Memory
TMA – Trap-based Memory Architecture
[email protected]
Dissertation Seminar
Nov 18, 2005
Vasaloppet
“Contention Problem in Sweden”
Traditional cross-country ski race
90 km …
[email protected]
Dissertation Seminar
85.6533
km to
go…
CS
Nov 18, 2005
Critical Section (CS) Cost
Spin Locks under Contention
Spin locks
Spin locks
with backoff
IF (more contention)
THEN less efficient CS …
“The more important the slower it runs…”
Amount of Contention
[email protected]
Dissertation Seminar
Nov 18, 2005
Queue-based Locks
CS Cost
Spin locks
Spin locks
with backoff
locks
IF (more contention)Queue-based
THEN constant CS cost …
Amount of Contention
[email protected]
Dissertation Seminar
Nov 18, 2005
CS Cost
This Dissertation
Spin locks
Spin locks
with backoff
IF (more contention)
THEN more efficient CS …
Queue-based locks
“The more important the faster
it runs…”
NUCA locks
Amount of Contention
[email protected]
Dissertation Seminar
Nov 18, 2005
NUCA Locks (Basic Idea)
1) Reduce Switch
traffic
- one CPU per node is testing…
Memory
$
$
P
P
…
2) ImproveMemory
lock handover
3) More efficient CS
$
P
Lock/Unlock
[email protected]
- local traffic is cheaper
$
$
P
P
…
Memory
$
$
$
P
P
P
Test
Test
Test
Test
Lock/Unlock
Dissertation Seminar
…
$
P
Test
Test
Test
Test
Test
Test
Test
Nov 18, 2005
The HBO Lock (the simplest HBO)
What do we need?
node_id
Creates
Compare&swap (CAS)
atomic operation
Communication
CAS(Lock_address, FREE, node_id)
lock-acquire:
Affinity
If the lock-value is in the state FREE:
• The node_id is CAS-ed into the lock location
Else: 2 cases
• The lock is “local” Spin with small backoff
• The lock is “remote” Spin with large backoff
Simple but fairly effective…
[email protected]
Dissertation Seminar
Nov 18, 2005
Performance Results
Realistic microbenchmark, 2-node WildFire, 28 CPUs
14
12
60
14
WF
50
Spin
MCS
HBO
10
9
Node Handoffs [%]
Iteration Time [seconds]
11
40
Fairness?
30
8
7
6
5
20
10
4
0
3
0
500
1000
1500
critical_work
[email protected]
2000
0
500
1000
1500
2000
critical_work
Dissertation Seminar
Nov 18, 2005
Fairness Study
Number of Finished Processors
Realistic microbenchmark, 2-node WildFire, 28 CPUs
t
28
26
24
22
20
18
16
14
12
10
8
6
4
2
0
Spin
MCS
HBO
0
[email protected]
5
Time [seconds]
Dissertation Seminar
10
15
Nov 18, 2005
Application Performance
28-processor runs
MCS
Spin EXP
Spin
Normalized Speedup
2.5
HBO
≈ 4x
2
1.5
1
0.5
[email protected]
Dissertation Seminar
ge
er
a
at
er
W
Av
-N
sq
nd
lre
Vo
ce
R
ay
t ra
io
si
ty
R
ad
FM
M
le
sk
y
C
ho
Ba
rn
es
0
Nov 18, 2005
Total Traffic: Raytrace
Local Transactions
Global Transactions
1.4
1.2
1
0.8
0.6
0.4
0.2
0
Spin
[email protected]
Spin EXP
Dissertation Seminar
MCS
HBO
Nov 18, 2005
HBO Locks inside Linux Kernel
Patch provided by Silicon Graphics, Inc.
Linux-IA64 kernel implementation, May 2005
Page-fault handler runs 3x faster
60 processors
[email protected]
Dissertation Seminar
Nov 18, 2005
Outline
NUCA Locks
DSZOOM – Software-based Shared Memory
TMA – Trap-based Memory Architecture
[email protected]
Dissertation Seminar
Nov 18, 2005
The DSZOOM Proposal
[email protected]
Dissertation Seminar
Nov 18, 2005
The DSZOOM Proposal
Run entire protocol in requesting-processor
No protocol agent communication!
Assumes user-level remote memory access
put, get, and atomics [ InfiniBand ]
Fine-grain memory protocols (e.g., 64 bytes)
Hardware-like memory models
[Shasta, Blizzard, Sirocco]
[email protected]
Dissertation Seminar
Nov 18, 2005
“Squeezing” Protocols into Binaries…
DSZOOM
Program
Original
Program
...
cmp
bne
nop
...
cmp
bne
nop
%g0, %l5
0x24431
ld
[%o1 + 64], %o0
ldd
clr
...
[%o0 + 16], %f4
%l5
%g0, %l5
0x24431
ld
mov
and
cmp
bne
nop
[%o1 + 64], %o0
255, %g6
%g6, %o0, %g6
%g6, 170
0x24450
ldd
clr
...
[%o0 + 16], %f4
%l5
Fast-path
Protocol
Code
Slow-path
Protocol
Code
(C-code)
Binary/Assembler level instrumentation
[email protected]
Dissertation Seminar
Nov 18, 2005
Write Permission Caching
Problem: store instrumentation relies on locking
More complex instrumentation
Solution: write permission cache (WPC)
Small and fast software-managed cache
Keeps write permissions
The WPC idea:
Exploit store locality
Dynamically reduce the number of memory references
in store checking code
[email protected]
Dissertation Seminar
Nov 18, 2005
Other “Features”
Two kinds of protocols
Invalidate
Update
Many optimizations
Instrumentation scheduling
Instrumentation batching
WPC-based write batching
WPC-based dirty-data filtering
Private-data filtering
# of WPC entries
Coherence unit size
[email protected]
Dissertation Seminar
(update and invalidate)
(invalidate)
(update)
(update)
(update)
(update and invalidate)
(update and invalidate)
Nov 18, 2005
Coherence Flags and Profiling
Coherence flags
Similar to optimization flags of compilers
Possible scenario:
gcc -dszoom-cl 128 -dszoom-inv –O3 my_app.c
Execution profiling
Similar to profile feedback of compilers
Helps finding appropriate coherence flag settings
Low overhead implementation in DSZOOM
• Less than 30 percent overhead
Works for both small and large input sets
[email protected]
Dissertation Seminar
Nov 18, 2005
DSZOOM Results
2-node WildFire, 16 CPUs
[email protected]
Dissertation Seminar
BEST
1.11x
ra
di
os
ity
ra
yt
ra
ce
w
at
er
-n
sq
w
at
er
-s
p
an
-n
c
oc
e
an
-c
oc
e
fm
m
ba
rn
es
ra
di
x
lu
-n
c
lu
-c
1.45x
PROFILED
ra
ge
inv-dwpc-64
av
e
inv-64
3.0
2.8
2.6
2.4
2.2
2.0
1.8
1.6
1.4
1.2
1.0
0.8
0.6
0.4
0.2
0.0
fft
Normalized Execution Time
HW-DSM
Nov 18, 2005
Outline
NUCA Locks
DSZOOM – Software-based Shared Memory
TMA – Trap-based Memory Architecture
[email protected]
Dissertation Seminar
Nov 18, 2005
Instrumentation Drawbacks
DSZOOM
Program
Original
Program
...
cmp
bne
nop
...
cmp
bne
nop
%g0, %l5
0x24431
ld
[%o1 + 64], %o0
ldd
clr
...
[%o0 + 16], %f4
%l5
•
•
ld
mov
and
cmp
bne
nop
[%o1 + 64], %o0
255, %g6
%g6, %o0, %g6
%g6, 170
0x24450
ldd
clr
...
[%o0 + 16], %f4
%l5
Binary transparency?
Run-time execution overhead
[email protected]
%g0, %l5
0x24431
Dissertation Seminar
Fast-path
Protocol
Code
Slow-path
Protocol
Code
(C-code)
Nov 18, 2005
Trap-Based Memory Architectures
Basic idea
Detect fine-grained coherence violations in hardware
Trigger a coherence trap when one occur
Maintain coherence by software protocols
No memory system modifications
Minimal processor modifications
Binary Transparency
No need to instrument binaries/applications
[email protected]
Dissertation Seminar
Nov 18, 2005
TMA Lite
Proof-of-concept Implementation
Load permission check
Hardware implementation of software check
• Predefined “magic-value” convention
Store permission check
Hardware WPC
Can be seen as a very small cache
Operates on virtual addresses
Accessed in parallel with the data TLB
[email protected]
Dissertation Seminar
Nov 18, 2005
TMA Lite Performance
[TMA: simulation study, 4 nodes | DSZOOM: 2-node WildFire]
HW-DSM
DSZOOM
DWPC
PROFILED
BEST
TMA
Normalized Execution Time
2.5
1.75x
2
1.01x
1.5
1
0.5
[email protected]
Dissertation Seminar
av
er
ag
e
at
er
-s
p
w
at
er
-n
sq
w
ra
di
x
lu
-n
c
lu
-c
fft
0
Nov 18, 2005
Topics not Presented
RH lock algorithm
Controlled (un)fairness
HBO_GT and HBO_GT_SD algorithms
Global throttling and starvation detection
DSZOOM implementation details
Instrumentation challenges; scheduling, batching, etc.
Bandwidth filtering techniques; dirty- & private-data
Innovative TMA simulation tricks
Low-level “good days” hacks
Reusing Simics checkpoints
[email protected]
Dissertation Seminar
Nov 18, 2005
Dissertation Seminar, 18/11 – 2005
Auditorium Minus, Museum Gustavianum
Software Techniques for
Distributed Shared Memory
Zoran Radovic
[email protected]
[email protected]
Dissertation Seminar
Nov 18, 2005