PowerPoint プレゼンテーション - University of Tokyo
Download
Report
Transcript PowerPoint プレゼンテーション - University of Tokyo
A guided wave approach to plane-to-plane optical
interconnects for multistage networks and
multiprocessor computers
2D folded perfect
shuffle permutation
Multistage hypercube computer
(2)
Fiber module
input
output
Processor arrays
Alvaro Cassinelli*, Makoto Naruse*,** and Masatoshi Ishikawa*
Univ. of Tokyo*, CRL**
Plan of the presentation
I. Multistage architecture for optical parallel computers
Reconfigurable multi-stage architecture
Hypercube and omega network examples.
II. Optical fiber-based interconnection module.
Why guided optics?
Module decomposition
III. Prototype fabrication and test
4x4 exchange prototype..
Transmittance, alignment tolerances
IV. Conclusion.
Present and future research directions
I. Multistage architectures for optical parallel computers
Hybrid optoelectronic
Data
flow
Interconnection
module
Photodetector
array
Interconnection
module
Interconnection
module
VCSEL array
Elementary
Processor Array
All optical
Data
flow
Interconnection/
switch module
Interconnection/
switch module
…
Interconnection/
switch module
…
A CReconfigurable
:
I.1
multi-stage architecture: principle
We will concentrate on network-based parallel computers (or “direct-connection machines”) rather
than on shared memory model (PRAM) as an efficient way to implement parallel computer
architectures. This choice is dictated by the fact that dealing with read/write conflicts in PRAM
machines is more related with control and routing, and we are primarily interested in topology and
communication primitives from a hardware point of view in our research –enhanced communication
primitives is what optical technology offers.
Optical technology offers enhanced parallel communication primitives
…of great benefit for network-based parallel computers
= distributed memory
shared memory
Static
Dynamic
Reconfigurable
interconnection
Pn
P1
controller
…switches inside
processors (local control)
…
…
Z
…
Pn
…
Fixed
interconnection
(X, Y, and Z)
…
Mem
Y
…
ULA
P2
…
X
…
control
…
P2
X
…
P1
Y
mux
Z
(X, Y or Z).
…switches outside processors
(local or global/external control possible)
A C :Dynamic architecture vs. static
I.2
Actually, in our former research, we studied
single-stage dynamic interconnections with
global control of the switch, using spatial light
In an n-degree
static
modulators (OCULAR
II).
topology, each processor
should have n distinct
optoelectronic I/O ports…
…
…
…
…
P1
interconnections
…
switches
…
processors
P2
…
…
…
Pn
Technologically challenging
Non reusable architecture
Bad scalability
Static networks can
be redesigned as
single-stage
P1
dynamic
networks…
Pn
…
P2
Feed-back loop
…processors, switches and
interconnections located in
distinct modules
Optimal use of electronic, optoelectronic and optics
Scalability, hardware reusability in other topologies
possible introduction of multiple stages…
I.3 The multi-stage paradigm
architecture can be “spanned” into
Stage 2
P1
P2
…
…
Pn
Cube Cycle
Tree
[computing]
Mesh
Pyramid
De Bruijn
P1
…
Pn
Delta
Omega
P2
…
S&I-1
P1
P2
Hypercube
Stage m
S&I-m
Stage 1
Multi-Stages
S&I-2
Single-Stage
Pn
Benes
Clos
[computing & networking]
Shuffle/exchange
Banyan
Simplicity (switches can be elemental 2x2 cross-bars)
The cost of
multiplying the
processors is paid
back as
Scalability / Reconfigurability for different topologies
Possibility of pipelining
Theoretical background: Multi-stage architectures have
been studied for decades in networking applications
A Coptical
:
I.2 The theoretically best
architecture (connectivity)
The linear architecture may be “sub-optimal” (Ozatkas)
when addressing thermal dissipation issues, but offers PD and VCSEL
Processor
much Photo-detector
easier “scalability”. Also, the “flat” optimal
flip-chip bonded
VCSEL
Elements (PE)
architecture,
willarray
work well with reflective holograms,
but to processor
array
(PD)
Array
array
would be much difficult to build using fiber arrays.
X
Optical (2D)
Data flow
connection
module
connection
module
…
MOAn(X) = A(n).I(n)… A(k).I(k)… A(1).I(1) (X)
Computation
made on PE
array
Optical shuffle of
data between PE
arrays
Matrix
representation of
computations on
the Multistage
Optical
Architecture
a) Free-Space reconfigurable interconnections
Optoelectronic
processing module
OCULAR-II
Elementary Processor Array
Photo-detector array
reconfigurable
interconnection
module
SLM-based
reconfigurable
interconnection
VCSEL array
reconfigurable
interconnection
module
reconfigurable
interconnection
module
Space-invariant interconnections – good/bad?
Free-space – alignment issues?
Multi-level CGH – good diffraction efficiency
Reconfiguration (“switch”) freq. – 100 Hz…
b) Fixed interconnections (hybrid opto-electronic)
OCULAR – III
Fixed interconnection modules...
Processor array in charge of the switching function…
Data
flow
Interconnection
module
Photodetector
array
Interconnection
module
Interconnection
module
…
VCSEL
array
Elementary
Processor Array
No lost of interconnection capacity if things are designed properly
Some examples: shuffle/exchange networks, Clos and Benes crossbars, etc…
A CTwo
:
I.3
well known examples:
- Ring, Mesh, and Hypercube are all classes of k-dimensional
Indirect Binary Cube (“multistage hypercube”)
nearest-neighbor
[ computingnetworks.
]
Binary
Hypercube…
-In the indirect binary hypercube network (as well as in the
Generalized Cube), we CAN NOT find the exchange
permutation E(k) at the end of stage k, we have to “wait” till
the end (the unshuffle is necessary…).
Y X
- The FFT isZan algorithm that is easily embedded in a
(2)
E(1)
P0= E(1)
hypercube W
topology. Moreover, the shuffle-exchange
“direct”
binary hypercube architecture can be used to demonstrate a
“pipelined” FFT algorithm very easily, because:
-1.,
(Omega)
network
FFT=(IBnC)
(4)
E(1)
(4)
E(1)
(4)
E(1)
(4)
(3)
E(1)
(4)
E(1)
-1(4)
feed-back
[ networking ]
E(1)
(IBnC)-1=
where
0000E(1) has been replaced by W(1). Of course,
0000
0001
0001
“Direct”
Binary
n-Cube
or
“generalized
Cube
network”.
0010
0010
0011
0100
0101
0110
0111
1000
1001
1010
1011
1100
1101
1110
1111
Output
Input
Self routing: “switches” are set
-The Omega network is also very useful for primitives of locally by packet address
parallel computing, like FFT algorithms!! Omega network is(destination – input)
0011
0100
0101
0110
0111
1000
1001
1010
1011
1100
1101
1110
1111
NOT full connection, it is full access BUT blocking. A non
It is full access, but not full
blocking network (rearrangeable, no “strict-non blocking”
connection.
nor “wide sense non-blocking”) is the BENES network.
Also useful on computing (FFT)…
CLOS is another which is strict-non blocking (but the
network is not constructed using 2x2 cross-bar switches).
II. Optical fiber-based plane-to-plane interconnection modules
(2)
Fiber module
input
…an optical “3D optical wiring” module
between 2D VLSI arrays.
output
II.1 Fiber-based interconnection blocks for multistage architectures.
• Inter-stage connection fixed and point-to-point: channels can be fibers.
• Fibers have better efficiency and just like free-space optics, no cross-talk in 3D.
• No space-invariance required.
• Precise and robust alignment possible.
• Theoretically more volume efficient than free-space equivalent!
Prototype Fiber module
(fibers and holders)
(2)
“integrated”
2D folded
perfect shuffle
permutation
module
• Maybe “hard” to build? Boring, but not a
fundamentally difficult - can be automated…
output
input
“Volume-consumption comparisons of
free-space and guided-wave optical
interconnections”, Y.Li and J. Popelek,
p.1815-1825, Appl.Opt. Vol 39, n.11, april
2000.
• Alignment of both output and input needed…
• Power dissipation may be a fundamental
limitation, but we are far from these limits…
…wave-guide arrays for fixed, point-to-point
and space variant interconnections are an
interesting alternative to free-space optics
A C :“Decomposition” of the interconnect into modules
II.2
In group-theoretic-based construction of MINs
(giving
symmetric
[ Problem
] networks), the most useful
permutations are the exchange, the shuffle, the
Many
networks
butterfly,
themultistage
bit reversal
and the use
shiftsimple, regular interconnections…
permutation.
The nature of the decomposition of the interconnect
into EITHER the column, row or the diagonal
dimensions may also reintroduce the use of lightefficient one dimensional, non pixilated, rapid
reconfigurable diffractive elements (such as acousto
However, when folded in a plane, these may materialize as non-regular, nonoptics).
scalable and non-reusable interconnection modules!
rows
columns
0
4
8
12
0
2
8
10
1
5
9
13
1
3
9
11
2
6
10
14
4
6
12
14
3
7
11
15
5
7
13
15
Scan map
Fractal map
[Solution ]
Because it may be possible to cascade fiber-based modules without too much
loss of light power, let’s “break” these into simple to fold modules.
“simple to fold” means:
1) Simple to implement by stacking planer wave-guide structures
vertical
horizontal
diagonal
Permutations are
decomposed “ad-hoc” into
their “row” and “column”
exclusive permutations
parts, plus some simple-tofold “link” permutation…
2) Or simple to implement using previously built modules (scalability)
Permutations are
decomposed “recursively”
The idea is to define permutation “constructors” that correspond to basic building
steps using PLC circuits (stacking, grouping modules).
Permutation
layer Pn/2
Ln/2 Pn/2
Permutation
module Pn
Rn/2 Pn/2
Z Pn
Vertical
replicator
Horizontal
replicator
“zoom”
constructor
Q Pn
“quadrant”
constructor
This decomposition methodology also applies to the switching stages (no other
thing that a set of possible permutations)
Let’s try that on the previous examples:
Indirect Binary n-Cube Network
…uses the butterfly (k)
and perfect shuffle (k)
permutations
P0= E(1)
-1
(3)
(2) E
E(1) (4) E(1) (4)
(1)
feed-back
0000
0001
0010
0011
0100
0101
0110
0111
1000
1001
1010
1011
1100
1101
1110
1111
E(1)
(4) E(1)
(4) E(1)
(4) E(1)
0000
0001
0010
0011
0100
0101
0110
0111
1000
1001
1010
1011
1100
1101
1110
1111
(Omega) network
Output
…uses only the
perfect shuffle (k)
permutation
Input
(4)
Example: shuffle and butterfly decomposition
Decomposition using constructors:
shuffle n(k)
{bn, … bk+1, bk, bk-1, … b2, b1}
n(k)
“ad-hoc”
n(k) = Ln/2 n/2; Rn/2 n/2 ; L
{bn, … bk+1, bk-1, bk-2, … b1, bk}
“ad-hoc”
butterfly n(k)
{bn, … bk+1, bk, bk-1, … b2, b1}
n(k)
{bn, … bk+1, b1, bk-1, … b2,bk}
…It is easy to see that the ad-hoc
folding of a “regular” permutation
needs a maximum of three
concatenated “stacked” modules
n(k) = Rn/2 n Ln/2 n/2; n/2 ; L
“recursive”
2p(2p) = Qp-1 T2 (1,2)
=
;
;
Folding the shuffle permutation
If k n/2, the shuffle “acts” over rows :
row(k)
(k)= row(k)
Can be built
by stacking
“slices”
If k > n/2, the shuffle can be written as:
(k) = row(n/2) .col(k-n/2).L
- where col is a column shuffle,
- and L is the “link” permutation.
Link
12
8
4=100
0
1=001
2
3
col(2)
AC:
Folding
a butterfly into a 4x4 array
REM: the modules can be built by stacking
layers,
thatthe
planar-optics
technology
If k ton/2,
butterfly (k)
“acts” over rows :
used. In particular, we can think again about
fan-in and fan-out channels… (cf. NHK
company.
row(2)
(2) = row(2)
If k > n/2, the butterfly (k) can be written as:
(k)= col (k-n/2).L.col(k-n/2)
- where col is a column butterfly,
- and L exchanges row and column LSB
Link
12
8
4=100
0
1=001
2
3
col(2)
…back to examples: network
shuffle
shuffle
shuffle
shuffle
row(2
90º
pair of PE implement
elemental exchange
switch
col(2
Processor arrays
(exchange switches and more)
L
I.3 Indirect Binary 4-Cube
PE array 1
(exchange)
PE array 2
(exchange)
PE array 3
(exchange)
(2)
PE array 4
(exchange)
(3)
(4)
Processor arrays
(exchange switches and more)
-1(4)
III. 4x4 prototype fiber module. Preliminary tests
Two holder prototypes: Zirconium, SiO2
Pitch: 250±5 m
Multimode graded index fibers: NA=0,21
(core 50m, cladding 126m)
Transmission loss: 3dB/km
Length: 30 cm
A C : Preliminary tests on a 4x4 prototype module
III.1
REM: the light coming from the non-addressed
[ Transmittance (one channel) ]
[ Interconnection
pattern
] default
channels
is mainly due
to some
functioning of the neighboring VCSELs which
(2)
emits LED light though
they are OFF!
45
40
38,45
Output
(CCD)
35
Transmittance (%)
Input
(VCSEL
854±4nm)
LED
regime
30
LASER
regime
25
20
15
10
5
0
6
7
8
9
10
11
12
9,5
VCSEL driving current (mA)
Max. transmittance 38,45% for I=9,5 mA
13
A C : Alignment tolerances (test performed on a single channel)
III.2
The differences on alignment tolerances are
probably due to the non-circular shape of
(2
Power
the VCSEL
mode.
output
input
meter
0.25
exit power (mW)
)
Horizontal excursion
0.2
x
0.15
VCSEL
ON
0.1
0.05
VCSEL array
X,Y, and Z
translation
stage
0
-105 -90 -75 -60 -45 -30 -15 0
15 30 45 60 75
X (microns)
No relay optics
between
VCSEL array
and fiber
module input
Alignment tolerances
(half peak power)
x 50 m
y 70 m
IV. Conclusion
AC:
Multi-function modules: the use of optical fiber modules fits well
[ Present
research
] for instance, one can imagine a module
with
the all optical
approach;
with several different interconnection patterns, but also other
“optical-functions”
like optical of
delay
lines:
Input/output alignment
modules
However, in all-optical networks the “switches” may be very fast
• Microlenses, Fibers with round ends.
(electro optical devices, not MEMS), because the delay time for
avoiding the drop
of ATM cells
?? for
a typical
Gigabit network!!!
• Modules
built is
from
fiber
bundles.
• Active alignment.
Demonstrator architectures using smart pixel arrays (2x2 or 4x4 electronic switches)
0
1
2
3
Optical
interconnection
[ Future research directions ]
Guided-wave interconnects can be “modulated” and integrated !
Multi-interconnection modules
• “Mixed” interconnections, and other optical functions
• Circuit switching for all optical networks
• Packet switching in a buffered architecture with globally controlled stages
Integrated plane-to-plane multistage paradigm
• using permutation “slices” for intra-chip massive, regular interconnections.
AC:
Multi-permutation
module
Rem: Dynamic alignment is
Interleaved
permutations
tightly
coupled with
dynamic into the same module: multi-permutation/switch module
reconfiguration of the
interconnect.
A small controlled mechanic or optical perturbation
Cf. Naruse’s presentation.
can produce a drastic change of the interconnection
pattern from input to output.
(…optical switches does not need to be “local” –i.e,2x2)
actuators
Use of MEMS technology?
inputs
“Normal” directional coupling between waveguides?
outputs
Transparent circuit switching by TDM interconnections
control
“all-optical” multistage
architecture
PE array
…optical switches does not
need to be “local” (2x2)
{ (1) , i}
{ (2), i}
{ (3), i}
bi-module
bi-module
bi-module
We are now building a demonstrator using mechanical displacement of modules
containing a by-pass interconnection and cube interconnections
“spanned” hypercube with
weak-communication
A new paradigm for packet switching in multistage networks
module control
module control
bi-module
{ (2), i}
bi-module
PE array
{ (1) , i}
PE array
PE array
input
output
module control
{ (3), i}
bi-module
…globally controlled exchange stages + Intermediate buffers
Selection method: alternate / Backpressure: on / mode: disablehop
4
4
0.9
0.8
3
0.7
3
0.6
2
0.5
2
0.4
1
0
0.3
1
0.2
64x64 Crossbar
64x64 MIN
64x64 GS-MIN
0.1
0
0
0.1
0.2
0
0.3
0.4
0.5
0.6
0.7
0.8
Input request probability (per unit time)
0.9
1
Length of buffers
Normalized Throughput (bandwidth)
1
Integrated multistage architecture?
waveguide “permutation slices”
WG
Normal
coupling
photonic
structure
- 3d IC integration of regular interconnected circuits
- a nice application for photonic bandgap coupling structures