DOI: 10.3390/rs18162701 ISSN: 2072-4292

Fine-Grained Urban Vegetation Segmentation Under Two Imaging Views Based on Scale-Aware Mixture of Experts and Scene-Specific Optimization

Yuhe Hu, Yujie Li, Nan Chen, Yuzhen Zhang, Yangle Jin, Yiqiu Chen, Jia Wang

High-precision urban vegetation mapping is essential for assessing carbon sink capacities, mitigating the urban heat island effect, and supporting sustainable development. Although deep learning and high-resolution remote sensing have advanced automated vegetation monitoring, existing models still face challenges when a common segmentation architecture is evaluated under different imaging geometries. In this study, Cityscapes and ISPRS Vaihingen are treated as two independent benchmarks representing perspective street-level imagery and orthographic aerial imagery, rather than as simultaneous cross-view inputs. “Background dominance” caused by perspective distortion and the “gridding artifacts” inherent in orthographic textures severely constrain segmentation accuracy across varying vegetation scales, particularly for small targets. To address these limitations, we propose a Scale-Aware Mixture of Experts (SA-MoE) architecture for fine-grained vegetation segmentation under two distinct imaging views, together with a scene-specific optimization strategy. The core SA-MoE framework consists of two main components. First, the spatial gating network uses a temperature polarization mechanism with τ = 0.5 to adjust the initial logit maps, sharpening expert-weight differences while preserving stable gradient propagation. Second, we use a heterogeneous expert group with five parallel branches: a pixel-level expert, three spatial experts with different dilation rates, and a global average-pooling expert. A dynamic pixel-level weighted fusion mechanism is then applied, decoupling feature extraction from receptive-field allocation. Furthermore, to address the heterogeneity of “hard samples” and “label noise” across the two benchmark settings, we introduce a scene-specific optimization strategy. Our findings show that the Focal-Dice (FD) loss is more suitable for perspective scenes with severe target imbalance and hard-to-classify vegetation targets, whereas the Cross-Entropy (CE) loss is more robust to boundary jitter in orthographic imagery. Comparative experiments on the Cityscapes (perspective view) and ISPRS Vaihingen (orthographic view) datasets reveal that SA-MoE achieves a highly competitive balance between computational efficiency and fine-grained segmentation, particularly in micro-target recall. Notably, the recall for extra-small (XS) scale targets in the aerial dataset improved by 3.21 percentage points compared to the second-best model. For the street-level dataset, our model achieved competitive global performance in terms of Overall Accuracy (OA), Precision, and F1-Score. However, we also observed a performance trade-off, where Transformer-based models maintained an advantage in preserving fine boundary details for these extra-small targets. In the routing analysis, we observed a pattern that we refer to as “receptive field inversion”, in which the model assigns lower weights to large-dilation experts for large canopy regions in orthophotos. We interpret this pattern as a plausible routing hypothesis. Overall, SA-MoE offers an efficient and adaptive solution for urban vegetation mapping under two imaging views.

More from our Archive