DOI: 10.3390/rs18152592 ISSN: 2072-4292

SPFMamba: A Mamba-Based Network with Semantic Prompt and Frequency-Adaptive Fusion for Remote Sensing Image Semantic Segmentation

Manlin Wang, Xifu Sun, Jiahang Liu, Yue Ni, Jian Cui, Ji Luan

Modern remote sensing images (RSIs) provide increasingly fine spatial detail, making pronounced scale variations and complex spatial distributions of land-cover classes more apparent and thereby increasing the difficulty of semantic segmentation. Recent remote sensing semantic segmentation methods therefore have increasingly adopted Transformer- and Mamba-based models to improve global contextual modeling. However, these models often suffer from intraclass inconsistency and interclass feature confusion, leading to fragmented object structures and imprecise boundary delineation. Accordingly, we develop SPFMamba, an architecture built around Mamba that combines semantic prompting with frequency-adaptive fusion, thereby improving global context modeling and fine-detail representation. During feature reconstruction, we propose a semantic prompt global–local Mamba (SPGLM) block to jointly model global semantic information and local spatial cues. Its parallel semantic prompt global and multidirectional local perception branches promote semantically coherent and spatially continuous feature distributions, thereby effectively preserving the structural integrity of ground objects. To further alleviate cross-level semantic discrepancies and strengthen the representation of small-scale targets, we design a high-frequency adaptive fusion module (HFAFM). It first refines high-frequency responses in shallow layers to retain small-scale object structures and boundary cues. Subsequently, deep semantic priors guide local cross-attention, allowing low-level spatial cues to be selectively integrated with deeper semantic representations. Evaluations across ISPRS Vaihingen, ISPRS Potsdam, and OpenEarthMap datasets show that SPFMamba delivers favorable segmentation accuracy with only 17.85 M parameters, while showing improved preservation of object structures and fine spatial details in qualitative comparisons.

More from our Archive