MSFusion: Multi-Scale Cross-Modal Fusion with Adaptive Attention for Multimodal Medical Image Fusion
Liu Wang, Yang Zhou, Wenjia Li, Lijuan ShiMultimodal medical image fusion integrates complementary information from heterogeneous imaging modalities to provide comprehensive visual support for clinical analysis. Most existing methods adopt an “encode–fuse–decode” paradigm that applies a single fusion rule only at the deepest network layer, often discarding shallow detail features and yielding blurred outputs with poor textural fidelity. To address this limitation, we propose MSFusion, a novel hierarchical framework that distributes adaptive fusion throughout the entire decoder stage. By leveraging skip connections to align decoder layers with corresponding encoder features, MSFusion enables full-scale integration of multi-resolution representations. The architecture employs a dual-branch convolutional encoder and introduces two core modules in the decoder: (1) the Multi-Scale Adaptive Fusion (MSAF) module, which addresses insufficient exploitation of cross-scale complementarity by dynamically weighting features via learnable attention, thereby balancing fine details and global semantics, and (2) the Multi-Scale Cross-Modal Cooperative Fusion (MSCMCF) module, which mitigates semantic misalignment through a cross-modal interactive attention mechanism that establishes robust inter-modality correspondences and promotes deep feature alignment. Additionally, a Vision RWKV (VRWKV) block is integrated to efficiently model both local and global spatial dependencies with linear computational complexity. Extensive experiments on public CT–MRI, PET–MRI, and SPECT–MRI datasets are evaluated using six standard metrics. On the CT–MRI benchmark, our method achieves MSE (↓) = 0.0423, CC = 0.8198, and SCD = 0.7023, outperforming state-of-the-art approaches. These results—combined with superior visual quality—demonstrate that MSFusion sets a new standard for accurate, detailed, and clinically meaningful multimodal image fusion.