An ensemble learning framework for protein stability prediction with enhanced recognition of stabilizing mutations
Yang Liu, Jian Zhang, Minghui LiAbstract
Accurately predicting mutation‐induced protein stability changes remains a central challenge in structural bioinformatics. Existing methods exhibit a strong bias toward destabilizing mutations, leading to limited performance for stabilizing mutations and constraining their utility in protein engineering. Here, we address this limitation through two complementary strategies: balanced dataset construction and integrative modeling. To mitigate the severe class imbalance in current stability datasets, we constructed undersampling‐based balanced datasets and further evaluated reverse‐mutation augmentation as a comparative strategy. Building on the rapid development of high‐performing predictors, we hypothesized that integrating their outputs could exploit complementary strengths and improve predictive accuracy. Accordingly, we developed three modeling frameworks, including models based on handcrafted features, models using embedding representations extracted from ProteinMPNN, and ensemble models integrating a diverse set of state‐of‐the‐art predictors. Across multiple independent test sets, ensemble models consistently outperformed individual approaches, with particularly pronounced gains in identifying stabilizing mutations. These findings demonstrate that combining undersampling‐based balanced data construction with systematic predictor integration provides an effective and practical strategy for achieving more balanced and accurate protein stability prediction, and offers a useful framework for identifying stabilizing mutations in protein engineering and related applications. StaMutAble is freely available at: