Reliable Machine Learning Screening of Adsorption Energies Is Better Assessed with Formula-Grouped Cross-Validation
Wenjie Wu, Mingling Yang, Ping Cheng, Yangning WangMachine learning (ML) models trained on bulk-crystal descriptors are increasingly used to prescreen catalysts by predicting adsorption energies, yet reported performances often rely on random K-fold cross-validation that permits the same bulk formula to appear in both training and test sets. We construct a reproducible benchmark that fuses 936 CatApp DFT adsorption energies with bulk descriptors from the Materials Project for H*, O*, and OH* on metal and alloy surfaces. We compare random K-fold cross-validation with GroupKFold grouped by parsed formula, the latter mimicking the realistic task of predicting adsorption on entirely new catalyst compositions. Under formula-grouped evaluation, random CV materially overestimates apparent generalization performance, with the largest and most robust effects for H* and OH* (protocol-inflation gaps up to approximately 0.8). The H* and OH* results are based on only 20 and 31 unique formulas, so their GroupKFold Spearman point estimates should be read as directional evidence rather than quantitative estimates. O* shows a smaller and statistically fragile protocol-inflation signal and, even where composition-plus-bulk features improve Random Forest and Ridge, the usable signal is best described as a very coarse pre-filter within a limited domain. Bulk descriptors are adsorbate-dependent: they improve O* prediction for Random Forest and Ridge, but degrade H* and OH*—a qualitative, directional observation given the small formula counts—whose binding is poorly captured by bulk crystal descriptors, consistent with the established view that it is governed by surface-localized electronic structure. These results outline a realistic performance boundary for bulk-to-surface ML in this benchmark: O* can be very coarsely prioritized from bulk descriptors within a limited domain, whereas H* and OH* are unlikely to be quantitatively predicted from bulk descriptors alone and would benefit from surface-aware models. We therefore recommend that bulk-to-surface adsorption-energy benchmarks report formula-grouped cross-validation alongside random cross-validation as a more robust and transparent practice.