DOI: 10.3390/rs18183219 ISSN: 2072-4292

SHALA: Sparse Hierarchical Alignment for Visible-Infrared Vehicle Re-Identification Under Aerial-Ground Viewpoint Gaps

Dong Liu, Xiaolin Zhao, Siyuan Zhao, Le Ru, Shuo Li, Pengfei Liu, Jincheng Bai

Visible-infrared vehicle re-identification across aerial-ground UAV platforms, where high-altitude and low-altitude views are separated by a large viewpoint gap, supports all-day traffic monitoring and cross-platform target association. However, most existing cross-modal re-identification methods are developed for single-platform camera networks with comparable viewpoint geometry across modalities, while cross-view re-identification is studied predominantly in the visible spectrum; the few works that combine the cross-modal dimension with large viewpoint or elevation gaps target pedestrians or elongated non-vehicle objects, and none has been developed for vehicle association in settings where modality discrepancy, aerial-ground viewpoint deformation, and sparse infrared texture co-occur. To address these limitations, we propose Sparse Hierarchical ALignment for Aerial-ground (SHALA), an end-to-end framework built on a frozen DINOv2 backbone. Modality-specific low-rank adapters and learnable view tokens provide modality adaptation and viewpoint compensation with only a small number of trainable parameters. Building on this, Top-K sparsification restricts each patch’s bidirectional cross-modal aggregation to its most relevant counterparts, which limits the influence of background patches and preserves the one-to-many correspondences induced by viewpoint compression. A parallel channel-response stability criterion anchors the foreground representation. On the real-world UCM-VeID benchmark, SHALA attains 41.8%/42.6% Rank-1 in the V2I/I2V directions, surpassing the strongest same-task baseline HWDNet by 2.2%/2.0%. On viewpoint-stratified CARLA pairs, Rank-1 degrades by only 7.2% as the pitch-angle gap widens, whereas representative baselines drop by 18.0–22.5%; the same method ordering reappears on real cross-platform Gold-Standard captures, where SHALA degrades more gracefully than both baselines, and real-data fine-tuning narrows the CARLA-to-UCM-VeID Rank-1 gap from 33.4 to 6.3 percentage points, with the CARLA held-out Rank-1 decreasing from 62.1% to 48.1% as real data is introduced.