DOI: 10.3390/rs18162700 ISSN: 2072-4292

DBCS-T: A Dual-Branch Cross-Attention Synergistic Transformer for Multimodal Image Fusion and Semantic Segmentation

Yiming Liu, Bo Gao, Xiao Yang, Hang Li, Weixing Yu, Huangrong Xu

Spectro-polarimetric imaging systems can simultaneously acquire spatial, spectral, and polarimetric information during remote sensing, yet the multimodal fusion data are often constrained in practical applications by insufficient exploitation of complementary information across different modalities. To address this issue, we propose a multimodal image fusion method based on a Dual-Branch Cross-Attention Synergistic Transformer (DBCS-T). In our method, three complementary feature components are independently extracted, i.e., a Characteristic Polarization Image (CPI), a Characteristic Spectral Image (CSI) and a Characteristic Intensity Image (CII). For CPI, it is derived from Angle of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) inputs via DBCS-T, which integrates a Cross-Channel Transposed Attention (CCTA) module for cross-modal interaction and a Multi-Scale Polarization Feature Adaptive Modulation (MPAM) module for local feature enhancement. CSI is obtained by leveraging maximum-divergence spectral band differences guided by prior spectral radiance curves. CII is computed from the Stokes parameter S0. These three components are then fused via Principal Component Analysis (PCA) into a Multimodal Fusion Image (MFI). To demonstrate the effectiveness of our method, an experiment was conducted on a scene that contains real vegetation, artificial foliage, and same-color metallic objects. The experimental results show that the proposed method achieves effective semantic segmentation of all three target categories. Furthermore, quantitative evaluation demonstrates that the MFI attains the lowest Kullback–Leibler (KL) and Jensen–Shannon (JS) Divergence values among all evaluated modalities, with image entropy exceeding that of individual source inputs. These results validate the complementarity of the extracted multimodal features, significantly enhance the interpretation performance for complex scenes, and demonstrate the broad application potential of the proposed fusion framework in remote sensing and multidimensional imaging.

More from our Archive