A Comparative Study of Machine Learning Approaches for Concrete Compressive Strength Prediction With SHAP and LIME Explanations
Saddam Hamada Ahneet Mohammed, Semaa Amer Ghanam, Qabas A. Hameed, Saif Saad Mohammed Khuder, Musaria Karim Mahmood, Muyasser Mohammed Jomaa’hAccurate prediction of the compressive strength of the concrete is crucial to optimize the mix design and construction quality and to save time and cost on experimental testing. In this study, three gradient boosting decision tree (GBDT) algorithms (extreme gradient boosting [XGBoost] light gradient boosting machine [LightGBM], category boosting [CatBoost]), three state‐of‐the‐art tabular deep learning (TDL) models (TabTransformer, neural ordinary differential equations [NeuralODEs], Tabular Deep Learning Meets Nearest Neighbors [TabR]) and a proposed stacking ensemble model for concrete compressive strength prediction are comprehensively compared. A set of 1030 actual concrete samples with eight input variables (age [days], cement [kg/m 3 ], Blast_Furnace_Slag [kg/m 3 ], Fly_Ash [kg/m 3 ], water [kg/m 3 ], superplasticizer [kg/m 3 ], fine aggregate [kg/m 3 ], and coarse aggregate [kg/m 3 ]) were used. A total of 12 input features were obtained after feature engineering (total binder, water‐to‐cement (W/C) ratio, water‐to‐binder (W/B) ratio, and aggregate‐to‐binder ratio). The data were split into 80% training and 20% testing sets and tenfold stratified cross‐validation used for model development and hyperparameter optimization. Six statistical metrics, RMSE, nRMSE, MAE, R 2 , MAPE, and max error, have been used to evaluate the performance of the models, while the Friedman statistical test, Nemenyi statistical test, and the Holm–Bonferroni corrected Wilcoxon statistical test have been applied to validate them. Taylor diagram analysis and SHapley Additive exPlanations (SHAP)–local interpretable model‐agnostic explanations (LIME) explainability techniques have also been employed to validate the models. Among the individual models, XGBoost achieved the best predictive accuracy, with RMSE = 3.9132 MPa, MAE = 2.4764 MPa, R 2 = 0.9387, and MAPE = 8.36%, followed closely by LightGBM and CatBoost. The proposed stacking ensemble, combining XGBoost, LightGBM, and CatBoost through ridge regression, achieved comparable predictive performance, with RMSE = 3.913 ± 0.858 MPa, MAE = 2.653 ± 0.401 MPa, R 2 = 0.939 ± 0.028, and the lowest maximum prediction error (17.26 ± 6.88 MPa), indicating improved robustness in extreme‐error cases. The statistical analysis revealed significant overall differences among the evaluated models ( p < 0.05), while the Taylor diagram and SHAP–LIME analyses provided complementary assessments of predictive fidelity and model interpretability. Overall, under the experimental conditions and dataset size considered in this study, the GBDT‐based models demonstrated stronger predictive performance than the evaluated tabular deep learning architectures. The proposed stacking ensemble achieved predictive accuracy comparable to XGBoost while exhibiting a lower maximum prediction error, suggesting improved robustness in extreme‐error cases. These findings support the proposed framework as an effective and interpretable approach for concrete compressive strength prediction and data‐driven decision support.