Transcript Poster
Fast Block Copy in DRAM
Motivation
Exploit the wide bandwidth within DRAM chip
Idea
Use DRAM refresh period to do block copy
Methodology
Add logic in DRAM
Extend the ISA of simulator
Conclusion
For 15740
The improvement depends on the system’s memory behavior
Fast Block Copy in Dram
1
Three Block Copy Modes
source
•Aligned Row Copy
For 15740
destination
•Unaligned Row Copy
Fast Block Copy in Dram
•Subrow Copy
2
DRAM Block Diagram
- Aligned Row Copy
RAS
Row
Row
address
address
latch
latch
10
1024
Row
Row
decoder
decoder
1024X1024
256 X256
cell
cellarray
array
row
1024
A0-A9
column
column
sense/write
sense/write
amps
amps
R/W
col
1024
CAS
Column
Column
address
address
latch
latch
For 15740
1024
WS
MUX
MUX
10
Column Latch
Column Latch
/ Decoder
/ Decoder
Fast Block Copy in Dram
Buffer
Buffer
Register
Register
3
Simulation
Simulator - SimpleScalar 2.0
Add new instruction – blkcp
DEFINST(BLKCP,
0x2e,
"blkcp",
"t,o(b)",
WrPort,
F_MEM|F_LOAD|F_STORE|F_DISP,
DCGPR(BS), DNA,
DGPR(RT), DGPR(BS), DNA,
({int index;
for (index=0; index<1024; index++)
WRITE_BYTE(READ_SIGNED_BYTE(GPR(BS)+OFS+index),
GPR(RT)+index);
}))
Rewrite memcpy() and bcopy() using blkcp
Rebuild benchmarks using new library routines
For 15740
Fast Block Copy in Dram
4
(a) Num ber of Instructions (IN)
1024
512
256
128
64
32
16
4
1024
0
512
0
256
10000000
5000000
128
15000000
20000000
10000000
64
30000000
32
25000000
20000000
16
50000000
40000000
8
30000000
4
60000000
8
Experiment 1 – Mem Copy
(b) Num ber of Mem ory References (MRN)
600000
The ideal effect of block copy
on memory system (Aligned row
400000
copy, x-axis: block size in blkcp)
1000000
800000
200000
1024
512
256
128
64
32
16
8
4
0
Best improvement achieved at
medial block sizes.
(c) Num ber of blkcp
For 15740
Fast Block Copy in Dram
5
Experiment 2 – File Read
(a) Num ber of Instructions (unaligned row copy)
1024
512
256
128
64
32
16
8
1024
512
256
128
64
32
16
8
4
4
60000000
50000000
40000000
30000000
20000000
10000000
0
60000000
50000000
40000000
30000000
20000000
10000000
0
(b) Number of Instructions (aligned row copy)
Performance improvement on file system
•Unaligned (Fig. a):
•The same effect as block copy intensive system
•Aligned (Fig. b):
•Best improvement achieved when block size = 4B & 8B
•Caused by different alignment of fread()’s internal buffer
For 15740
Fast Block Copy in Dram
6
Experiment 3 – Perl
(SPECint95)
1024
512
256
128
64
32
16
8
4
16000000
14000000
12000000
10000000
8000000
6000000
4000000
2000000
0
Num ber of Instructions
Triangle points represent simulation errors
# of memcpy() wrt different block sizes
Limited performance improvement:
caused by limited memory block copy
For 15740
Fast Block Copy in Dram
7
Conclusion
Good for block copy intensive system
Limitations
Do not support multiple memory banks
Hardware cost and overhead
Only user mode behavior, only two routines changed
Future work
For 15740
Simulate both kernel and user modes
Use more realistic benchmarks
Compare with data prefetching and non-blocking cache
Fast Block Copy in Dram
8