DOI: 10.3390/rs18193271 ISSN: 2072-4292

MHBA-TransUNet: Shallow–Deep Collaborative RGB–DSM Fusion with Hybrid Bidirectional Attention for High-Resolution Remote Sensing Semantic Segmentation

Tanghong Yuan, Wu Le, Ming Lv, Zhenhong Jia, Jiajia Wang, Gang Zhou, Sensen Song, Tingji Han

High-resolution remote sensing semantic segmentation using RGB imagery and digital surface models (DSM) remains challenging because of insufficient shallow cross-modal fusion and inadequate coordination between global semantics and local spatial information. To address these issues, we propose MHBA-TransUNet, a hybrid encoder–decoder network for RGB–DSM semantic segmentation. In the encoder, a Shallow Cross-modal Feature Fusion Module (SCFFM) progressively integrates complementary RGB and DSM features through dynamic spatial recalibration, while a lightweight dual-stream Transformer with a self-attention–cross-attention–self-attention (SA–CA–SA) configuration models intra-modal context and bidirectional cross-modal interactions. During decoding, a Global–Local Alignment Module (GLAM) combines Block-Distance Linear Attention (BDLA), feature-wise modulation, and channel reweighting to coordinate deep semantic features with shallow spatial details. Experiments on the ISPRS Vaihingen, ISPRS Potsdam, and US3D datasets achieve mIoU scores of 85.15%, 88.01%, and 85.49%, respectively, outperforming representative state-of-the-art methods. Ablation studies further verify the effectiveness and complementarity of SCFFM and GLAM, while the US3D results demonstrate stable performance across diverse urban scenes.