DOI: 10.1093/bioinformatics/btag710 ISSN: 1367-4811

Genome-wide-scale prediction of compound−protein interactions using foundation and language models based on three-dimensional structures of compounds and proteins

Yuga Moriyama, Yuki Matsukiyo, Tsuyoshi Kimura, Hiyori Yamamoto, Yuko Sakajiri, Tomokazu Shibata, Ryusuke Sawada, Yoshihiro Yamanishi

Abstract

Motivation

The identification of compound–protein interactions (CPIs) is crucial in the early stages of drug discovery. However, machine-learning (ML)-based methods based on one- and two-dimensional representations cannot capture important geometric information on the binding sites of CPIs, which limits their predictive accuracy.

Results

We propose DESTIN (dual structure-guided estimation of substance–target interaction networks), an ML framework for CPI prediction on a genome-wide scale from the three-dimensional (3D) structures of compounds and proteins. We constructed compound feature vectors that spatially reflect stable 3D conformations from a quantum–chemical perspective by fine-tuning a 3D structure-based foundation model. Furthermore, we constructed protein feature vectors leveraging sequence-based structure representations and protein language models, even for proteins that do not have 3D structures determined experimentally. CPI prediction was performed using an integrative ML model on 3D structure-aware feature vectors for compounds and proteins. The effectiveness of the proposed model was validated via rigorous benchmark tests on practical scenarios. DESTIN exhibited good predictive performance, even for structurally complex and pharmacologically important membrane proteins. These results show that DESTIN is promising for various pharmaceutical applications.

Availability

All source code and data supporting this study are publicly available on GitHub and archived on Zenodo (https://doi.org/10.5281/zenodo.21759701).

Supplementary information

Supplementary data are available at Bioinformatics online.