DOI: 10.3390/su18168458 ISSN: 2071-1050

Interpretable Machine Learning for Monthly Mean Air Temperature Modeling Under Correlated Meteorological Predictors: A Single-Station Case Study in Zonguldak, Türkiye

Rukiye Uzun Arslan, İrem Şenyer Yapici, Berna Aksoy

Reliable modelling of monthly air temperature is relevant to station-scale climate assessment and the evaluation of meteorological data-driven models. However, station-scale monthly meteorological datasets often contain correlated and partially redundant predictors because thermal, moisture, precipitation, wind, and seasonal variables are jointly controlled by atmospheric and seasonal forcing. This study conducts an integrated comparative analysis of established regression and machine learning models for monthly mean air temperature modelling in Zonguldak, a humid coastal province in the Western Black Sea Region of Türkiye. Monthly meteorological observations from 2000 to 2022 were used to evaluate eight primary regression and machine-learning models: Partial Least Squares regression, Ridge, Lasso, ElasticNet, Support Vector Regression, Random Forest, Gradient Boosting, and Extreme Gradient Boosting. Ordinary Least Squares (OLS) and Huber regression were additionally included as reference models. The analysis retained the original meteorological predictors and jointly evaluated predictive accuracy, model stability, ablation sensitivity, and model-specific predictor relevance. Reduced-predictor and seasonality-only scenarios were examined to distinguish direct thermal reconstruction from broader climatological predictability. Model performance was assessed using repeated nested cross-validation, bootstrap summaries of performance variability, supplementary rolling-origin validation, and Wilcoxon signed-rank tests with Holm correction. Although the full-predictor models achieved high predictive accuracy, this performance largely reflected the direct thermal information contained in minimum and maximum air temperature. When these thermal predictors were excluded, MAE increased to approximately 1.13–1.22 °C and R2 decreased to approximately 0.93–0.94. The seasonality-only scenario yielded MAE values of approximately 1.27–1.34 °C and R2 values of approximately 0.92, indicating that the annual cycle accounted for a substantial proportion of monthly temperature predictability. The additional non-thermal meteorological predictors provided only limited improvement beyond the strong seasonal baseline. Overall, model performance depended on the predictor information available, and no single model family showed a consistent advantage across the evaluated scenarios. These findings highlight the importance of considering predictive accuracy together with model stability and predictor dependence in data-limited station-scale temperature modelling.

More from our Archive