DOI: 10.3390/en19184426 ISSN: 1996-1073

Day-Ahead XGBoost Forecasting of Aggregated Residential Load: Accuracy and SHAP Ranking Agreement Across Experimental Configurations

Piotr Szeląg, Tomasz Popławski, Michał Adamusiński

This study assessed how the selected seasonal test period, training-window strategy, and hyperparameter selection were associated with differences in XGBoost day-ahead forecast accuracy and interpretation for approximately 300 G11-tariff households in Poland. Sixteen configurations combined four 31-day periods, sliding or expanding training windows, and shared (H1) or window-specific (H2) hyperparameters. The model used 19 temporal, meteorological, and calendar features; meteorological predictors for training, validation, and testing were archived numerical weather prediction (NWP) forecasts from the same operational forecasting system, available before the forecasted day. Accuracy was evaluated using mean absolute error (MAE), root mean square error (RMSE), and mean absolute percentage error (MAPE); paired comparisons used the Diebold–Mariano test with the Harvey–Leybourne–Newbold correction and Holm adjustment. Global SHAP (SHapley Additive exPlanations) rankings were compared using Spearman’s coefficient. MAE ranged from 7.79 to 14.98 kWh, and all configurations had lower MAE, RMSE, and MAPE than both persistence benchmarks. After Holm correction, no training-window strategy showed a statistically supported advantage; H2 was supported for the autumn expanding-window comparison, whereas the summer result depended on the variance estimator. The 24 h consumption lag ranked first in every configuration, and mean rank agreement across 120 pairs was 0.897. Accuracy varied across periods and configurations, whereas feature hierarchy remained highly consistent.