GCF-Net: Stage-Aligned Optical–Elevation Fusion for Aerial Remote Sensing Semantic Segmentation
Yifan Yu, Song Deng, Yang Yang, Fan MinOptical–elevation data fusion is widely used in aerial remote sensing semantic segmentation, as optical imagery provides rich spectral and textural information, while DSM or DEM data offer complementary elevation-related structural cues. However, effective fusion remains challenging because optical and elevation representations may exhibit cross-modal structural inconsistency, frequency–spatial response imbalance, and decoder-stage structural attenuation. To address these challenges, we propose GCF-Net, a stage-aligned optical–elevation fusion network that matches different cross-modal processing objectives to the evolving representation states of the encoder–decoder pipeline. A Structure-Guided Cross-Modal Correction Module first performs structure-conditioned correction of modality-specific features before fusion. A Frequency–Spatial Cross-Modal Fusion Module then constructs joint representations through bounded cross-modal magnitude conditioning, frequency-to-spatial reconstruction, and spatial recalibration. During decoding, a Geometry-Aware Cross-Scale Refinement Module reintroduces elevation-derived structural guidance into multiscale fused features. Experiments on ISPRS Vaihingen, ISPRS Potsdam, and MMHunan yield mIoU scores of 72.37%, 75.38%, and 52.23%, respectively, achieving the highest mIoU among the evaluated unimodal, multimodal, and SAM-based methods under the unified protocol. Ablation and replacement experiments verify the complementary roles of the three stage-specific components, while sensitivity, elevation perturbation, and complexity analyses indicate architectural flexibility, tolerance to moderate elevation degradation, and a balanced accuracy–efficiency trade-off.