AE‐VMUNet: Attention‐Enhanced Vision Mamba U‐Net With Multi‐Scale Feature Refinement for Prostate Tumor Segmentation in T2‐Weighted MRI
Xueting Wei, Yuehui Liao, Feng Gao, Ruipeng Li, Haote Chen, Panfei Li, Xiaobo Lai, Jingyu Zhu, Yasheng HuangABSTRACT
Accurate segmentation of prostate tumors on T2‐weighted magnetic resonance imaging (MRI) scans is essential for diagnosis and treatment planning. However, this remains challenging due to the heterogeneous appearance of tumors, their indistinct boundaries, and the frequent occurrence of small lesions. We propose AE‐VMUNet, an attention‐enhanced Vision Mamba U‐Net that performs stage‐wise feature refinement within a state‐space encoder‐decoder framework. The model integrates the following: (i) an Adaptive State‐space Enhanced Feature (ASF) block in the encoder, which couples long‐range dependency modelling with channel‐wise recalibration; (ii) an Attention‐Gate‐based Skip Connection Fusion (SCF) module, which performs spatially selective, multi‐scale aggregation; and (iii) an Efficient State‐space Guided (ESG) block in the decoder, which preserves global context while enabling detail‐aware reconstruction. Experiments on the public PI‐CAI 2022 dataset and an independent, private clinical cohort demonstrate that the AE‐VMUNet model achieves Dice scores of 59.51% and 48.83% respectively, with HD95 values of 14.21 and 18.96 mm. These results outperform those of representative CNN‐, attention‐ and Mamba‐based segmentation baselines. Further evaluations on PROMISE12 and ACDC demonstrate strong cross‐dataset and cross‐anatomy generalization, achieving Dice scores of 92.19% and 91.64%, respectively. Overall, these results suggest that combining state‐space modelling with lightweight attention mechanisms can lead to precise and robust medical image segmentation.