DOI: 10.1021/acs.jmedchem.6c00249 ISSN: 0022-2623

Interpretable Prediction of Ligand–Protein Binding without Protein Structural Information

Ananthan Sadagopan, Anurag Sodhi, William J. Gibson

Abstract

Ligand–protein binding prediction remains a central challenge, yet the contribution of ligand-side information to performance is unclear. We combined pretrained molecular embeddings with TabPFNv2 to build per-target classifiers without protein features. Across 159 BindingDB targets, models assigned higher probabilities to annotated binders and achieved >10-fold enrichment at the top 1% for 42 targets and >50-fold enrichment for three. Fragment- and atom-level interpretability analyses recovered established pharmacophores and nominated concise target-associated substructures. In a BRD9 DNA-encoded library screen, the model distinguished hits from nonhits from the same experiment (AUC = 0.913) and recovered the 2-pyridone chemotype. Supporting analyses separated carbonic anhydrase actives from matched DUD-E decoys, recovered primary and off-targets for compounds in DepMap, and guided the synthesis of a structurally simplified compound that measurably inhibited ACC2 ATPase activity. These results establish ligand-only models as interpretable screening tools and motivate their use as a baseline for assessing the added value of protein representations.

More from our Archive