Hierarchical Spatial Discrimination for UAV Geo-Localization Under Continuous-Coverage Partial Matching
Qiang Li, Huawei Liu, Jingzhi Zhang, Baoqing Li, Jianpo LiuCross-view geo-localization enables GNSS-denied UAV navigation by matching a drone image against a geo-tagged satellite database, a key capability for remote sensing. Existing methods perform well on landmark benchmarks but struggle in continuous-coverage scenarios, where a drone image only partially overlaps surrounding tiles and many works sidestep the issue by discarding partial pairs. This setting requires hierarchical spatial discrimination: the model must first separate distant regions and then distinguish nearby tiles using fine-grained spatial layout. We build this hierarchy into three pipeline stages. At the feature level, intermediate Vision Transformer (ViT) layers with Global Average Pooling provide an effective spatially aware representation at no extra parameter cost. At the training level, Dynamic Curriculum Mining (DCM), our primary training contribution, progressively shifts in-batch negative mining from coarse geographic discrimination toward fine-grained visual confusions. At the inference level, the Spatial Overlap Estimation Module (SOEM) estimates drone–satellite overlap from patch-token matching and re-ranks the top candidates. On GTA-UAV cross-area, our method achieves 71.75% Recall@1 and 222.0 m Dis@1; on DenseUAV, it reaches 96.22% Recall@1. In zero-shot cross-dataset evaluation on the real-world UAV-VisLoc dataset, the GTA-UAV-trained model achieves 35.39% Recall@1 without target-domain adaptation, providing additional evidence of transfer to real UAV imagery. Rather than relying on heavy auxiliary modules, the framework deliberately keeps additional architectural overhead minimal: layer truncation reduces backbone inference cost by about 16%, while SOEM introduces only 0.08 M additional parameters.