Transcript Poster

Fast Block Copy in DRAM
 Motivation

Exploit the wide bandwidth within DRAM chip
 Idea

Use DRAM refresh period to do block copy
 Methodology


Add logic in DRAM
Extend the ISA of simulator
 Conclusion

For 15740
The improvement depends on the system’s memory behavior
Fast Block Copy in Dram
1
Three Block Copy Modes
source
•Aligned Row Copy
For 15740
destination
•Unaligned Row Copy
Fast Block Copy in Dram
•Subrow Copy
2
DRAM Block Diagram
- Aligned Row Copy
RAS
Row
Row
address
address
latch
latch
10
1024
Row
Row
decoder
decoder
1024X1024
256 X256
cell
cellarray
array
row
1024
A0-A9
column
column
sense/write
sense/write
amps
amps
R/W
col
1024
CAS
Column
Column
address
address
latch
latch
For 15740
1024
WS
MUX
MUX
10
Column Latch
Column Latch
/ Decoder
/ Decoder
Fast Block Copy in Dram
Buffer
Buffer
Register
Register
3
Simulation
 Simulator - SimpleScalar 2.0
 Add new instruction – blkcp
DEFINST(BLKCP,
0x2e,
"blkcp",
"t,o(b)",
WrPort,
F_MEM|F_LOAD|F_STORE|F_DISP,
DCGPR(BS), DNA,
DGPR(RT), DGPR(BS), DNA,
({int index;
for (index=0; index<1024; index++)
WRITE_BYTE(READ_SIGNED_BYTE(GPR(BS)+OFS+index),
GPR(RT)+index);
}))
 Rewrite memcpy() and bcopy() using blkcp
 Rebuild benchmarks using new library routines
For 15740
Fast Block Copy in Dram
4
(a) Num ber of Instructions (IN)
1024
512
256
128
64
32
16
4
1024
0
512
0
256
10000000
5000000
128
15000000
20000000
10000000
64
30000000
32
25000000
20000000
16
50000000
40000000
8
30000000
4
60000000
8
Experiment 1 – Mem Copy
(b) Num ber of Mem ory References (MRN)
600000
The ideal effect of block copy
on memory system (Aligned row
400000
copy, x-axis: block size in blkcp)
1000000
800000
200000
1024
512
256
128
64
32
16
8
4
0
Best improvement achieved at
medial block sizes.
(c) Num ber of blkcp
For 15740
Fast Block Copy in Dram
5
Experiment 2 – File Read
(a) Num ber of Instructions (unaligned row copy)
1024
512
256
128
64
32
16
8
1024
512
256
128
64
32
16
8
4
4
60000000
50000000
40000000
30000000
20000000
10000000
0
60000000
50000000
40000000
30000000
20000000
10000000
0
(b) Number of Instructions (aligned row copy)
Performance improvement on file system
•Unaligned (Fig. a):
•The same effect as block copy intensive system
•Aligned (Fig. b):
•Best improvement achieved when block size = 4B & 8B
•Caused by different alignment of fread()’s internal buffer
For 15740
Fast Block Copy in Dram
6
Experiment 3 – Perl
(SPECint95)
1024
512
256
128
64
32
16
8
4
16000000
14000000
12000000
10000000
8000000
6000000
4000000
2000000
0
Num ber of Instructions
Triangle points represent simulation errors
# of memcpy() wrt different block sizes
Limited performance improvement:
caused by limited memory block copy
For 15740
Fast Block Copy in Dram
7
Conclusion
 Good for block copy intensive system
 Limitations



Do not support multiple memory banks
Hardware cost and overhead
Only user mode behavior, only two routines changed
 Future work



For 15740
Simulate both kernel and user modes
Use more realistic benchmarks
Compare with data prefetching and non-blocking cache
Fast Block Copy in Dram
8