Explainable Multi-Label Machine Learning Framework for Patient-Specific Antibiotic Recommendation
Saman I. Othman, Rebaz Hamza Salih, Kamal Al-Barznji, Karzan M. Abdullah, Muhamed Aydin Abbas, Shayma Ali Hussein, Blnd Azad Ismail, Mohammed Awat Ali, Ahmed Abdulrazzaq Bapir, Christer Janson, Aras Bradosty, Kardo I. Nuradin, Shukur Wasman SmailRapid selection of appropriate antimicrobial therapy is important for improving patient outcomes and supporting antimicrobial stewardship, particularly in the context of increasing antimicrobial resistance (AMR). Conventional antimicrobial susceptibility testing (AST), although essential for clinical decision-making, may not provide actionable susceptibility information immediately. This study develops an explainable multi-label machine learning framework for patient-specific antibiotic recommendation based on susceptibility prediction and ranked decision support using routinely collected clinical microbiology data. The framework integrates leakage-controlled preprocessing, microbiological and demographic feature representation, multi-label susceptibility encoding, One-vs-Rest ensemble learning, probability-based antibiotic ranking, statistical evaluation, temporal validation, and SHapley Additive exPlanations (SHAP). The dataset comprised 234 clinical records, of which 196 bacterial/other records were retained for the primary analysis after excluding 38 fungal records. A chronological partition produced 156 development records and 40 temporally held-out test records. Thirty-three antibiotic susceptibility labels were retained based on development-set availability. XGBoost, Random Forest, and LightGBM were evaluated using five-fold out-of-fold (OOF) validation. Random Forest was selected for the final recommendation and explainability analyses based on its overall performance, achieving a Micro-F1 of 0.4204, Macro-F1 of 0.3003, AUROC of 0.7069, AUPRC of 0.3979, and Precision@5 of 0.3962. A leakage-safe local antibiogram was additionally evaluated as a population-level ranking baseline, achieving a Precision@5 of 0.3077. Feature ablation showed that the combination of organism and age produced the highest OOF Micro-F1 (0.4635) and AUROC (0.7186), while the addition of gender and specimen type improved selected ranking measures but did not consistently improve aggregate classification performance. On the temporally held-out test cohort, Random Forest achieved a Micro-F1 of 0.4267, AUROC of 0.7447, AUPRC of 0.4892, and Precision@5 of 0.4350. SHAP analysis was completed for all 33 antibiotic-specific classifiers, providing global and antibiotic-level explanations of model behavior. The framework provides an interpretable approach for ranking potentially susceptible antibiotics at the patient level and is intended as clinical decision support rather than an autonomous prescribing system. Further external and prospective multicentre validation, including dedicated evaluation of challenging cases such as pan-drug resistance, is required before clinical deployment.