Compressive Strength Prediction of Self-Compacting Concrete with Recycled Coarse Aggregate Using Machine Learning: Robust Multi-Split Evaluation and Data-Leakage Analysis of a Stacking Ensemble
Nenad Kojić, Bojan MiloševićReliable prediction of the compressive strength of self-compacting concrete with recycled coarse aggregate (SCRCAC) from mixture composition supports more rational mix design and fewer experimental tests. Using the benchmark dataset of the reference study (603 mixtures, eight input variables), this work re-examines machine-learning prediction of this property with an emphasis on honest evaluation rather than on a new model. A stacking ensemble of three gradient-boosting models (XGBoost, LightGBM, CatBoost) and an extremely randomized trees model, combined through a ridge meta-learner, is used as a representative model and compared with the four machine-learning models of the reference study (Random Forest, Extra Trees, XGBoost, LightGBM), the recent single-booster model of Abood et al., and the reference artificial neural network. Reported as the mean over 25 repeated 70/30 splits, the ensemble reaches R2 = 0.793 ± 0.038 and RMSE = 6.21 ± 0.48 MPa, above all four reference models (R2 = 0.7249–0.7635) and significantly, though only marginally, above a tuned single XGBoost. The central contribution is the evaluation itself. Because the dataset contains repeated identical compositions, a leakage-free protocol lowers the R2 of every model to between 0.60 and 0.71, showing that the values of about 0.81–0.87 usually reported are inflated by duplicate-composition leakage, and leave-one-source-out evaluation lowers it further to about 0.14. Mutual-information and partial-dependence analyses identify cement as the dominant predictor, with water acting mainly through a nonlinear dependence. Robust, leakage-aware evaluation, rather than model architecture, emerges as the key to credible strength prediction on this benchmark.