Protein-based resistance in-silico model (PRISM-TB) for rapid prediction of drug-resistant tuberculosis using machine learning
Theja K.V., Ram Shankar Barai, Uma Devi Ranganathan, Swati Joshi, Achint Chaudhary, Agniva Das, Sukhdev MishraBackground and objectives
Tuberculosis (TB) caused by Mycobacterium tuberculosis (MTB), remains a priority health challenge with multidrug-resistant (MDR) strains threatening control efforts. Conventional diagnostic methods are limited by high cost, prolonged diagnostic time and incomplete mutation coverage. This study presents PRISM-TB (Protein-based resistance in-silico model for TB), a machine learning-based framework for predicting resistance to first- and second-line anti-TB drugs using protein sequence-derived features.
Methods
Protein sequences of 11 resistance-associated Mycobacterium tuberculosis genes were retrieved from NCBI and curated using literature-confirmed resistance mutations to generate a labelled dataset of 1,369 sequences across six anti-TB drugs. Sequence-derived numerical features were extracted and subjected to standardised preprocessing prior to model development. Six supervised machine-learning algorithms were trained and optimised using Bayesian hyperparameter tuning with stratified cross-validation, followed by model performance evaluation. External validation was performed using 461 mutations from the WHO Catalogue of Mutations in M. tuberculosis complex.
Results
Ensemble classifiers consistently outperformed linear and probabilistic models in predicting drug resistance from protein sequence derived features. Extra Trees achieved the highest overall performance (F1-score = 0.942, accuracy = 0.941), while Random Forest demonstrated superior class discrimination (ROC-AUC = 0.973) using only 10 features. On external validation, Random Forest correctly identified 351 of 374 WHO-graded resistant mutations (recall = 0.939, accuracy = 0.818). Feature importance and SHapely Additive exPlanations (SHAP) analyses identified physicochemical descriptors related to charge, hydrophobicity, polarity, and sequence transitions as the primary drivers of resistance prediction.
Interpretation and conclusions
This study highlights the potential of integrating protein-level mutation data with interpretable machine learning models for rapid, in-silico prediction of MDR-TB, offering a cost-effective approach to resistance screening.