Crossed Cubes

In-memory computing

In-Memory Computing for Machine Learning and Deep Learning

Wide-ranging review of in-memory computing for machine-learning accelerators: emerging-memory devices, matrix-vector primitives, digital/analog/mixed implementations, evaluation metrics, and device non-idealities.

N. Lepri, A. Glukhov, L. Cattaneo, M. Farronato, P. Mannocci, and D. Ielmini. IEEE Journal of the Electron Devices Society, 11, 587–601, 2023.

Read the open-access article

Parallel algorithms

Programming the Twisted-Cube Architectures

Develops programming methods for the multiply-twisted cube, including distributed optimal routing, SIMD algorithms, and the emulation of hypercube algorithms. It frames a lower-diameter recursive network as a usable parallel-computing architecture, not merely a graph construction.

K. Efe. Proceedings of the 9th International Conference on Distributed Computing Systems, 1989.

Read the article PDF (external copy)

Introduction to Parallel Algorithms and Architectures: Arrays, Trees, Hypercubes

A foundational treatment of parallel-algorithm design across arrays, trees, and hypercubes, including communication primitives, prefix computations, routing, sorting, and matrix algorithms.

F. T. Leighton. Morgan Kaufmann, 1992.

Read the publisher record

Foundational work on collective communication, routing, scans, and matrix algorithms will be added here. These sources provide the algorithmic context for the communication studies on this site.

Dataflow and accelerator architectures

TPU 8t and TPU 8i Technical Deep Dive

Google's first-party overview of its eighth-generation TPU systems. It distinguishes the training-oriented TPU 8t, its MXU and 3D-torus scale-up network, from the TPU 8i system, which adds a Collectives Acceleration Engine and the Boardfly topology for low-latency serving.

D. Gupta and S. Mugazambi. Google Cloud Blog, April 22, 2026.

Read the article

Sources on systolic arrays, dataflow architectures, and matrix multiplication will be added here. They offer useful comparison points for computation organized around regular local movement.

Optical interconnects

Space-Time-Coded High-Speed Reconfigurable Card-to-Card Free-Space Optical Interconnects

An experimental 2 x 2, 10-Gb/s card-to-card free-space optical interconnect for data centers and high-performance computing. Particularly relevant to board-facing optical links, it also measures the effect of air turbulence on link range and sensitivity.

K. Wang, A. Nirmalathas, C. Lim, K. Alameh, and E. Skafidas. Journal of Optical Communications and Networking, 9(2), A189-A197, 2017.

Read the article

Present Status and Future Needs of Free-Space Optical Interconnects

Board-to-board and chip-to-chip review of free-space optical interconnects. It identifies density, distance-bandwidth product, power, crosstalk, packaging, alignment, and thermal stability as central design questions.

S. Esener and P. Marchand. Materials Science in Semiconductor Processing, 3(5-6), 433-435, 2000.

Read the article

Optical Interconnects for Extreme Scale Computing Systems

System-level review of optical interconnects for large HPC systems, including bandwidth-per-FLOP pressure, optical switching, power, and cost tradeoffs.

S. Rumley, M. Bahadori, R. Polster, S. D. Hammond, D. M. Calhoun, K. Wen, A. Rodrigues, and K. Bergman. Parallel Computing, 64, 65-80, 2017.

Read the article

Interconnection networks

The Crossed Cube Architecture for Parallel Computation

Introduces the crossed cube as a lower-diameter alternative to the hypercube, develops its distributed routing and SIMD algorithms, and examines its ability to emulate hypercube computation.

K. Efe. IEEE Transactions on Parallel and Distributed Systems, 3(5), 513–524, 1992.

Read the article

References on hypercubes, topology-aware routing, and network structure will be added here.