Optimizing GPU-aware halo exchange for computations on structured grids
Johannes Pekkilä, Touko PuroHalo exchange is the task of communicating the boundaries of a subdivided computational grid with neighboring processes in distributed computations. With graphics processors augmenting the world’s fastest supercomputers, data movement across heterogeneous memory hierarchies must be carefully orchestrated to leverage these accelerators at scale. In this work, we implement and optimize packing, inter-process communication, domain decomposition, and topology-aware process assignment integral to efficient halo exchange on heterogeneous systems using GPU-aware MPI. Our test cases include benchmarks of the halo exchange on one- to five-dimensional grids on both AMD and Nvidia graphics processors. We find that in multidimensional halo exchange, custom packing kernels provide up to two orders of magnitude speedup compared to MPICH and OpenMPI implementations based on MPI datatypes. Furthermore, hierarchical topology-aware process assignment provides more consistent performance across varying grid sizes on heterogeneous systems compared to simpler methods. We apply our methods improve the strong scaling of magnetohydrodynamics simulations, demonstrating 1.58× speedup on high process counts in parallel computation and communication. Our results highlight the importance of kernel fusion for alleviating small kernel overheads on graphics processors even if it prohibits pipelining packing with communication.