Enzymatic Reaction Feasibility Classification Using Machine Learning Methods
Xin Wang, Hongyan Yin, Yekai Shen, Yushan Zhu, Igor V. Tetko, Aixia YanAbstract
With the advancement of computer-aided retrobiosynthesis, numerous biosynthetic pathways have been predicted, exceeding the capacity of experimental validation. Effective classifiers are needed to identify feasible reactions. In this study, we collected 75,864 feasible reactions from biocatalysis databases and generated an equal number of infeasible reactions based on reaction rules. Following atom mapping focused on reaction centers and systematic reaction preprocessing, three datasets for training were constructed: the “stereo” dataset, which retained reaction stereochemical information; the “non-stereo” dataset, which was a stereochemistry-agnostic version of the “stereo” dataset; and the “mixed” dataset, which comprised both. We established a total of 22 individual enzymatic reaction feasibility classification models, which include: eXtreme Gradient Boosting (XGBoost) and Deep Neural Network (DNN) models utilizing Reaction Fingerprints (RXNFP), Differential Reaction Fingerprint (DRFP), and our constructed Combined ECFP4 Reaction Fingerprints (c_ECFP4) for reaction representation, and Transformer models and fine-tuned ChemBERTa-77M-MLM (ChemMLM) models using reaction SMILES strings as the direct input. The results indicate that models utilizing the c_ECFP4 representation achieved the highest predictive performance, which effectively captured underlying enzymatic reaction mechanisms. Among them, Model 1A-M (based on XGBoost and “mixed” dataset) was identified as the optimal individual model, achieving Matthews Correlation Coefficient (MCC) values of 0.865 and 0.853 and Area Under Curve (AUC) values of 0.981 and 0.980 on the “stereo” and “non-stereo” test sets, respectively. Furthermore, a consensus model enzymatic reaction feasibility classification (ERFC) integrating four reaction representations further improved predictive performance, achieving MCC values of 0.894 and 0.886 and an AUC of 0.986 on both test sets. Moreover, both models (Model 1A-M and ERFC) successfully validated a five-step biosynthetic pathway, demonstrating higher prediction accuracy than the previously reported DeepRFC and DORA-XGB models in identifying feasible reactions. All data, the individual Model 1A-M, and the consensus model ERFC are openly available, offering a reliable and flexible framework for predicting enzymatic reaction feasibility in the presence or absence of stereochemical information.