An Interpretable Multi-Objective Machine Learning Framework for In Silico Prioritization of Anti-Staphylococcus aureus Antimicrobial Peptides
Jianguo Xu, Donghua Yang, Qingyong Zheng, Tengfei Li, Yating Cui, Jinhui TianBackground: Staphylococcus aureus, including methicillin-resistant lineages, is a leading cause of device- and catheter-related infection, and rising resistance motivates the search for antimicrobial peptides (AMPs) with strong anti-staphylococcal activity and low host toxicity. Machine learning can prioritize candidate peptides. However, the literature-derived AMP datasets are prone to homology-driven optimism, and computational studies frequently overstate their translational reach. Methods: We curated 4007 deduplicated S. aureus-active AMP records and 582 binary-labeled hemolysis records. Each peptide was encoded with a transparent 538-dimensional physicochemical and compositional feature vector. Five classifiers and five regressors were evaluated for four endpoints (potency classification, log10 MIC regression, hemolysis classification, normalized hemolytic index) under both conventional random 5-fold cross-validation and homology-aware cross-validation, in which sequences were clustered by 3-mer similarity and whole clusters were confined to single folds. Class imbalance was handled by class weighting. Model behavior was interpreted with SHAP and Fisher-exact k-mer enrichment, and candidates were ranked by a multi-objective score that combines the independently trained heads. Results: Under homology-aware validation, performance was lower than under random splitting, as expected. Potency classification reached an AUROC of about 0.71 (Random Forest), compared with 0.797 under random cross-validation. Hemolysis classification remained strong at AUROC 0.90 (95% CI 0.88 to 0.93), which indicates that its high accuracy is not a homology leakage artifact. MIC regression was modest (homology-aware R2 0.17, Spearman ρ 0.39) and is therefore treated only as a rank-ordering signal. SHAP and k-mer analyses recovered interpretable structure–activity relationships. Net positive charge and amphipathicity drove potency, whereas bulk hydrophobicity drove hemolysis. Applying the pipeline to a generated pool prioritized 20 candidates. Nearest-neighbor analysis shows that these are close optimized variants of known potent scaffolds, with a median identity of 95% to a known peptide, rather than novel sequences. Conclusions: We present an interpretable, honestly benchmarked multi-objective pipeline that optimizes known anti-S. aureus AMP scaffolds toward lower predicted hemolysis. The prioritized peptides are computational hypotheses for future synthesis and experimental testing. Their low predicted hemolysis reflects a selection criterion rather than validated safety, and cross-species selectivity was not assessed.