Mol-ME: Enhancing Molecular Property Prediction via Multi-Modal Alignment Learning and Ensemble Methods
Baoren Huang, Mu Chen, Junjie Luo, Lei Wang, Wencai Ye, Ke WangAbstract
Molecular deep learning plays an important role in addressing challenging molecular property prediction tasks. However, labeled molecular data remain scarce, and the majority of existing studies predominantly employ single-modal methods. Most single-modal models face limitations in simultaneously capturing molecular topological features and modeling long-range dependencies in sequences. In this study, we propose a novel multimodal alignment framework for joint modeling of molecular graphs and sequences, called Mol-ME. The framework incorporates a data augmentation strategy to enhance model performance under limited labeling conditions. Mol-ME comprises four core modules. The first module consists of dual encoders that generate graph-based and sequence-based molecular representations, which are then aligned through contrastive learning. The second module, a gated cross-modal fusion network, enables fine-grained integration of these representations by leveraging both the cross-attention mechanism and the gating mechanism. The third module is a motif-aware feature extractor that captures latent relationships among molecular substructures. The final module employs ensemble learning to predict on extracted representations, which captures complex nonlinear relationships and compensates for the modeling limitations of single shallow networks. Experimental results on 9 benchmark data sets demonstrate that Mol-ME consistently outperforms all baseline methods, achieving new state-of-the-art (SOTA) performance in molecular property prediction.