Federated Spectral Regularization for Convergence Acceleration: A Random Matrix Theory Perspective
Shengyu Cai, Jianchao BaiFederated learning enables privacy-preserving distributed training but suffers from client drift and slow convergence under statistical data heterogeneity. Most existing federated optimization methods address client drift via parameter-space constraints or aggregation-level corrections, while fewer works directly shape the gradient covariance spectral structure of the optimization landscape. This paper analyzes the convergence problem from a spectral perspective, revealing that non-IID data causes spectral diffusion in the gradient covariance matrix and degrades convergence. Guided by random matrix theory, we propose federated spectral regularization (Fed-SR), a computationally efficient method that indirectly constrains spectral spread via gradient norm regularization. Although computing the regularizer gradient requires Hessian vector products, our optimized auto-differentiation implementation avoids storing full Hessian matrices and restricts extra computational overhead to a negligible level. Experiments on CIFAR-10, CIFAR-100, and other benchmarks show that Fed-SR outperforms baselines including FedAvg, FedProx, and SCAFFOLD in non-IID scenarios, reducing communication rounds and improving accuracy and stability. Ablation studies, spectral analysis, and controlled spectral feature manipulation experiments provide consistent empirical evidence showing a strong empirical association between the “spectral concentration” effect and performance gains, offering mechanistic interpretability consistent with our proposed theoretical framework within the tested experimental settings.