DOI: 10.1137/25m1731794 ISSN: 1064-8275

Accelerating Neural Network-Based Regression and Classification Tasks through Empirical Kernels

Saad Qadeer, Andrew Engel, Amanda Howard, Adam Tsou, Max Vargas, Panos Stinis, Tony Chiang

Abstract.

The impressive performance of deep neural networks (DNNs) on a variety of learning tasks has spurred much investigation into improving their training and characterizing the trained functions. Recent work has shown the equivalence of DNNs in the infinite width limit and kernel machines relying on the neural tangent kernel (NTK) at initialization. These results suggest, and experimental evidence corroborates, that kernel machines relying on empirical kernels extracted from trained DNNs can act as surrogates for trained finite-width DNNs. The high computational cost of assembling the NTK, however, makes this approach infeasible in practice. In the current work, we study the performance of the conjugate kernel (CK), an efficient approximation to the NTK. For smooth function and logistic regression, we show that the CK performance is only marginally worse than that of the NTK and, in certain cases, much more superior. In particular, we establish bounds for the test losses, verify them with numerical tests, and identify the regularity of the kernel as the key determinant of performance. We also determine regimes where both kernel machines trained on features extracted from an underlying DNN are demonstrably superior to the latter and use this to suggest a recipe for accelerating DNN performance inexpensively. We present a demonstration of this on foundation models by comparing their performance on a classification task using a conventional technique and our prescription. We also show how our approach can be used to improve physics-informed operator network training as well as convolutional neural network training for vision classification tasks.

More from our Archive