A Robust Model Evaluation Process for Early-Stage Cooling Load Prediction of Buildings
Yaren Aydın, Ümit Işıkdağ, Sinan Melih Nigdeli, Gebrail Bekdaş, Wook-Won Kim, Zong Woo GeemIn the construction industry, a large portion of energy is spent on heating and cooling, which both increases costs and contributes to resource depletion. The aim of the study was to provide and evaluate a robust ML model evaluation process for early design stage cooling load prediction of buildings. For this purpose, 18 different machine learning models were evaluated using a Nested Cross-Validation approach consisting of 50 outer fold and 50 inner Optuna trials, along with hyperparameter optimization. To avoid model selection being dependent on small decimal differences, paired model comparisons, effect sizes, Holm-corrected statistical tests, and the 1-SE economy rule were applied over the same outer folds. As a result of the analysis, Categorical Boosting (CatBoost) was selected as the final model, and within the Nested-CV framework, R2 = 0.8275 ± 0.0072, RMSE = 1.6792 ± 0.0210 kWh, MAE = 1.4356 ± 0.0233 kWh, and MAPE = 0.0521 ± 0.0009 were obtained. Model interpretability analyses showed that the variables Ambient Temperature, Solar Radiation, and Heat Reflective Treatment had the highest permutation importance values. Residual analyses revealed that the model exhibited low systematic bias, but the residual variance was dependent on the estimate value, and the residuals deviated from a normal distribution. This study provides a framework that evaluates not only the prediction performance but also model selection, generalization stability, interpretability, and residual behavior together. The findings demonstrate that CatBoost is a strong option for cooling load prediction in this simulation-based dataset. However, validation of the obtained results with real building data and different climatic conditions is considered an important requirement for future studies in terms of evaluating the external validity of the model.