DOI: 10.1093/bioinformatics/btag711 ISSN: 1367-4811

OrchAlign: Orchestrated Multimodal Alignment for Gene Expression Prediction

Li Liu, Guipeng Xv, Shujie Liu, Yeyun Gong, Chen Lin

Abstract

Motivation

Gene expression prediction benefits from integrating DNA sequences and epigenomic signals. Existing approaches typically combine these modalities using simple operations such as concatenation or summation, without explicitly modeling fine-grained token-level cross-modal interactions. Establishing precise cross-modal correspondences remains challenging due to (1) inter-modal discrepancies, and (2) intra-modal structural preservation.

Results

To address these challenges, we propose OrchAlign, an orchestrated multimodal alignment framework that establishes fine-grained cross-modal alignment while preserving intra-modal structure. Our approach first disentangles each modality into shared and specific representations. We then perform focused cross-modal alignment on the shared subspace within a key regulatory region, producing aligned representations with token-level correspondences. Moreover, we enforce consistency between relational structures of the aligned shared representations and their corresponding modality-specific counterparts, enhancing intra-modal structural coherence. Finally, the aligned shared representations and modality-specific representations are passed to the BiMamba backbone for further multimodal fusion. Our experimental results show that OrchAlign achieves state-of-the-art performance in gene expression prediction.

Availability and Implementation

All resources are available at https://github.com/XMUDM/OrchAlign

Contact and Supplementary Information

Correspondence should be addressed to Chen Lin (chenlin@xmu.edu.cn). Supplementary information is available online.