ISAR-Mamba: A Dual-Stream Gated Mamba with Hierarchical Spatial Summaries for ISAR Image Captioning
Yonghua He, Aoxiang Pan, Yonggang Li, Jiahao Wang, Wei Qu, Weigang Zhu, Wenhang Ji, Guodian Tang, Junyi LvInverse Synthetic Aperture Radar (ISAR) imagery plays a crucial role in space situational awareness, yet its interpretation remains largely confined to tasks such as classification and segmentation, lacking the ability to generate detailed natural language captions. To address this limitation, this paper proposes ISARCap 1.0, the first dataset specifically designed for ISAR image captioning, covering 38 classes of space targets and comprising 102,965 simulated images and 514,825 image–text pairs. On this basis, this paper proposes ISAR-Mamba, a dual-stream gated state space model for ISAR image captioning. The model introduces a Dual-Stream Asymmetric Encoder (DSAE), which performs complementary patch partitioning and scanning along the range and azimuth dimensions, respectively, to accommodate the physical dimensional differences of ISAR images. Meanwhile, a Sparsity-Aware Gating Mechanism (SAGM) is introduced, which jointly suppresses the interference of empty patches on state updates by leveraging scattering energy priors and a learnable scoring module. Furthermore, a Hierarchical Spatial Summarization (HSS) strategy is proposed to extract structured summary tokens from different depths and spatial regions of the encoder, thereby enhancing visual information perception during text sequence generation. Experimental results on the ISARCap 1.0 dataset indicate that ISAR-Mamba outperforms existing image captioning methods on multiple metrics, suggesting the effectiveness of the proposed method and the usability of the dataset.