Bonsai: Efficient and Optimal Automatic Tensor Rematerialization for Memory-Constrained DNN Training
Dat Nguyen, Vasudha Devarakonda, Anxiao Jiang, Khanh NguyenGPU memory is increasingly the primary bottleneck in scaling deep neural network (DNN) training, where the activation tensors footprint of a model may exceed the memory capacity. Tensor recomputation is a powerful technique that trades additional computation for reduced peak memory usage. However, existing approaches face a fundamental tension between performance optimality and computational scalability. On the one hand, solvers leverage Integer Linear Programming (ILP) to provide mathematically optimal solutions but suffer from the combinatorial explosion of the search space and thus become intractable for modern DNN models. On the other hand, heuristics-based approaches achieve scalability but sacrifice optimality altogether, resulting in suboptimal execution schedules.
The root cause of these inefficiencies in the state of the art is the mismatch in abstraction. This paper introduces Bonsai, a framework that tackles this scalability-granularity tension. At the heart of Bonsai is a novel abstraction of operator segmentation that breaks the computation graph into flexible, variable-sized units to enable a lightweight yet effective segment-based ILP formulation. By having segments, Bonsai collapses the search space and prunes redundant solutions that stall existing solvers. This abstraction enables Bonsai to maintain a holistic view of the entire model, ensuring that no optimization opportunity is lost while reducing the number of decision variables by orders of magnitude. The evaluation across a diverse set of DNN architectures and models demonstrates that Bonsai scales to real-world models, is up to 10.13× lower solver cost than state-of-the-art ILP solvers, and delivers up to 11