ERβ-Score: An Interpretable Machine Learning-Based Scoring Function and Web Server for Estrogen Receptor β-Guided Drug Discovery in Triple-Negative Breast Cancer
Abbas Khan, Muhammad Ammar Zahid, Walid Kouidri, Osama Aboubakr Mohamed, Ahmed Mohammad Gharaibeh, Ladun Ibrahim Mohamed, Amani Anwar Al-Mansori, Mohamed Haitham Elsayed, Anwar Mohammad, Ameera Al-Jabiry, Mohanad Shkoor, Raed M. Al-Zoubi, Abdelali AgouniTriple-negative breast cancer (TNBC) is the most clinically aggressive subtype of breast cancer, characterized by the absence of targetable hormone receptors and HER2 amplification, significantly constraining treatment choices. Estrogen Receptor Beta (ERβ) has emerged as a biologically relevant yet underutilized target in TNBC, with its re-expression linked to tumor suppression and improved prognosis, prompting the development of selective ERβ modulators as a precision therapeutic approach. We introduce ERβ-Score, an interpretable machine learning scoring system developed using a curated dataset of 1699 ERβ bioactive chemicals obtained from ChEMBL, characterized by 39 physicochemical and three-dimensional molecular descriptors. After implementing scaffold-disjoint train/test partitioning to avert structural data leakage, a Gradient Boosting Classifier, fine-tuned through Bayesian hyperparameter optimization, attained in five-fold cross-validation a Precision–Recall AUC (Area Under the Curve) of 0.891, a ROC-AUC (Receiver Operating Characteristic) of 0.888, a Matthews Correlation Coefficient of 0.664, an F1-score of 0.838, and a balanced accuracy of 0.831; on the scaffold-disjoint hold-out test set it attained a Precision–Recall AUC of 0.905, a ROC-AUC of 0.864, and a Matthews Correlation Coefficient of 0.578, indicating strong and balanced discrimination between active and inactive ERβ modulators. We note explicitly that this scaffold-disjoint hold-out constitutes internal validation, since it derives from the same curated ChEMBL workflow used for model development, and it is therefore reported throughout as scaffold-disjoint internal validation rather than as independent external validation. The applicability domain boundaries were established using a k-nearest-neighbor Tanimoto-similarity method with ECFP4 (Extended-Connectivity Fingerprint with a Diameter of 4) fingerprints, offering a quantitative confidence metric that identifies structurally new molecules beyond the model’s reliable prediction range. External validation against independent Tox21 ERβ bioassay data confirmed genuine, statistically significant predictive signal (ROC-AUC = 0.71) while revealing reduced sensitivity for structurally novel active compounds. The model was subsequently used for extensive virtual screening of natural product and drug-like compound libraries, with prioritized candidates undergoing structure-based molecular docking against the ERβ co-crystal structure (PDB: 7XWQ) using Smina, facilitating a comprehensive evaluation of hits based on both ligand and structural properties. To enhance accessibility, the complete pipeline was implemented as an open-access interactive web application utilizing Streamlit, enabling researchers to input any SMILES string and obtain, in real time, an activity prediction with a probability score, applicability domain classification, Lipinski drug-likeness assessment, interactive three-dimensional visualization of protein–ligand interactions, and on-demand docking within the ERβ active site.