A Comprehensive Review of Multimodal Medical Image Fusion: Techniques, Evaluation, and Future Directions
Pengquan Han, Cuiyin Liu, Cuiwei WangIn recent years, multimodal medical image fusion (MMIF) has attracted significant attention due to its ability to integrate complementary information from different medical imaging modalities and provide more comprehensive information for clinical analysis. By combining anatomical information from modalities such as computed tomography (CT) and magnetic resonance imaging (MRI) with functional information from positron emission tomography (PET) and single-photon emission computed tomography (SPECT), MMIF can improve image quality and support subsequent tasks such as disease analysis, lesion detection, image segmentation, and treatment planning. This review provides a comprehensive overview of MMIF from theoretical and technical perspectives. First, commonly used medical imaging modalities and publicly available medical image databases are summarized and compared. Subsequently, the general workflow, fusion levels, and quality requirements of MMIF are introduced. Representative fusion techniques are then systematically reviewed, including spatial-domain methods, transform-domain methods, sparse representation-based methods, deep learning-based methods, hybrid methods, and emerging Mamba-based approaches. In addition, commonly used image fusion quality assessment metrics are analyzed, and the reported quantitative performance of representative MMIF methods is compared and discussed. Finally, current challenges and future development trends of MMIF are presented, including robustness, clinical translation, and emerging multimodal learning paradigms.