DOI: 10.3390/sym18081389 ISSN: 2073-8994

Cluster Resource Load Prediction Method Based on Temporal-Feature Attention and Dynamic Stacking

Qiaoyan Zhang, Kaijun Wu, Chenshuai Bai

Accurate prediction of computing-resource workloads is important for capacity planning, overload warning, and intelligent system management. Large-scale computing systems exhibit complex temporal fluctuations, sudden variations, and multi-resource coupling characteristics, making accurate workload prediction challenging. To address these challenges, this paper proposes a Hybrid Temporal-Feature Attention enhanced Multi-task Stacking model (HTAM-Stack) for multivariate cluster workload forecasting. First, a Temporal-Feature Hybrid Attention (TFHA) module is designed to jointly capture temporal dependencies and cross-resource feature interactions, enabling adaptive extraction of critical temporal patterns and important resource characteristics. Second, a Multi-Task Learning (MTL) framework is introduced to simultaneously predict CPU and Memory workloads by exploiting the correlations among heterogeneous resource variables. Furthermore, a Dynamic Stacking (DS) mechanism is developed to adaptively adjust the contributions of heterogeneous base learners through a weight generation network, and a Residual Corrector (RC) is incorporated to further enhance prediction robustness. Extensive experiments conducted on two widely used public cluster workload datasets, including Google Cluster Trace and Alibaba Cluster Trace, demonstrate that HTAM-Stack achieves competitive prediction performance under complex and dynamic workload conditions. The proposed model achieves MAE values of 0.0012 and 0.0010, RMSE values of 0.0031 and 0.0027, MAPE values of 0.20% and 0.16%, and R2 values of 0.9715 and 0.9782 on the two datasets, respectively. Moreover, HTAM-Stack requires only 3.20 ms inference time with 8.60 M parameters, achieving a favorable balance between prediction accuracy and computational efficiency. The results validate the general effectiveness of the proposed framework on public cluster workload benchmarks rather than its direct applicability to railway IT systems. Because no representative railway workload dataset was available, railway IT is discussed only as a potential application context that requires domain-specific validation.

More from our Archive