Wafer-Scale and Multi-GPU Systems for AI

Research questions
- How should address translation work across a wafer-scale GPU?
- How can page prefetching and memory techniques reduce the cost of data movement?
- How should network traffic use interconnects with non-uniform bandwidth?
- How do electrical and optical interconnect choices affect wafer-scale and multi-GPU systems?
Publications
2026
GPGPU 2026
RIPPLE: Ring-based Page Prefetching and Layered Translation for Wafer-Scale GPUs
2025
2022
GPU Architecture
2026
HPCA 2026
QuCo: Efficient and Flexible Hardware-Driven Automatic Configuration of Tile Transfers in GPUs