Mini-Batch Pruning Techniques in Stochastic Gradient Descent Methods
Roberto Cavoretto, Alessandra De Rossi, Enrico ScibiliaMini-batch Stochastic Gradient Descent (SGD) is a standard optimization method for training neural networks and other large-scale parametric models. We study three loss-based mini-batch pruning strategies—Perceptron-Inspired SGD (PISGD), Pruning SGD (PSGD), and No-Release PSGD (NRPSGD)—that reduce the number of per-sample gradient evaluations by deprioritizing low-loss samples during training. The methods are evaluated in Python on two problems: polynomial function approximation and MNIST handwritten-digit classification with a one-hidden-layer neural network. We examine the sensitivity to pruning hyperparameters and compare the three strategies with standard mini-batch SGD in terms of training time and optimization or predictive performance. In the reported MNIST experiments, the pruning methods show competitive or improved time–accuracy and time–loss behavior relative to standard mini-batch SGD, with NRPSGD obtaining the most favorable results in the longer runs. These findings provide proof-of-concept evidence restricted to the experimental settings considered. They indicate that loss-based mini-batch pruning can reduce gradient-computation cost without degrading solution quality in the tested settings. Because the pruning mechanism leaves the parameter-update rule unchanged, it may also be combined with adaptive stochastic optimizers such as RMSProp and Adam.