SLUG: Vision-Based UAV Self-Localization in GNSS-Denied Urban Environments
Xuxu Qi, Enhui Zheng, Beicheng Li, Chen Tian, Xuehao HuangIn complex environments such as cities and canyons, cross-view matching of UAV and satellite imagery serves as a complementary approach to self-localization in the absence of GNSS signals. To address the limitations of existing datasets—namely their limited variety of scenes and low annotation accuracy—we have constructed the Cross-View, Multi-Altitude, Multi-Scene UAV Localization Dataset (CMHUAV). This dataset covers nine adjacent, non-overlapping subregions in Hangzhou and comprises 8625 precisely matched pairs of drone and satellite images, providing high-precision supervision for cross-view feature learning. Furthermore, we designed a Transformer-based global feature enhancement localization model (SLUG) equipped with a lightweight Adaptive Part Weighting (APW) module, which learns part-level scalar importance weights to dynamically recalibrate the contribution of each annular region. By integrating adaptive part-level weighting with a joint supervision strategy, we significantly enhance SLUG’s spatial discrimination capability. We also propose the Height-aware Normalized Distance Matching (HNDM@K) metric, which unifies retrieval ranking quality and geospatial bias within a single framework, enabling joint evaluation of elevation-adaptive localization accuracy and retrieval performance. Experiments demonstrate that, on the CMHUAV dataset, SLUG achieves a recall of 93.53% and a mean precision of 93.9%; under the HNDM@1 metric, the localization scores at all three altitude levels exceed 94.8%. These results demonstrate that the proposed method, evaluated via offline visual retrieval, exhibits strong potential as a vision-based localization module for UAVs under GNSS disruption. Nevertheless, closed-loop flight tests and system-level localization performance still require further validation.