DOI: 10.3390/ph19101534 ISSN: 1424-8247

A Cluster-Guided Screening Framework for Prioritizing Natural Product Candidates Against HIV-1 from the KNApSAcK Database

Muhammad Alqaaf, Md. Abdullah Al Mamun, Ahmad Kamal Nasution, A S M Nazrul Islam, Retno Supriyanti, Naoaki Ono, Shigehiko Kanaya, Md. Altaf-Ul-Amin

Background: Plant-derived natural products are a productive antiviral scaffold source, yet secondary metabolite libraries remain underexplored against HIV-1 amid extensive target-structure redundancy. This study presents a cluster-guided framework for HIV-1 inhibitor prioritization from the KNApSAcK database. Methods: A total of 64,166 SMs were converted to SMILES and queried against BindingDB to identify reported HIV-1 integrase, protease, and reverse-transcriptase associations. A total of 295 HIV-1 protein sequences were aligned and partitioned using DPClusSBO; representative structures (6VDK, 1MUI, 6ELI) per cluster were docked against cluster-mapped SMs using SMINA. Prioritized SMs were evaluated with SwissADME and benchmarked against ChEMBL HIV-1 inhibitors. Results: Clustering resolved three non-overlapping groups (integrase, n = 135; protease, n = 117; reverse transcriptase, n = 42), mapping 285, 124, and 410 SMs, respectively. Predicted docking scores ranged from −18.91 to −4.72 kcal/mol (integrase), −26.54 to −7.87 kcal/mol (protease), and −26.83 to −4.59 kcal/mol (reverse transcriptase). ADME-prioritized reverse-transcriptase compounds scored more favorably than the matched NNRTI reference set (p < 0.001), requiring experimental confirmation of binding affinity. DUD-E enrichment validation showed strong discriminative validity for integrase and protease (ROC-AUC 0.85, 0.78) but not reverse transcriptase (ROC-AUC 0.53), consistent with weaker pose reproduction for the latter. Conclusions: The framework reduced target redundancy and computationally prioritized natural product candidates for HIV-1 as hypothesis-generating predictions requiring experimental validation.