DOI: 10.3390/rs18183238 ISSN: 2072-4292

Motion-Guided Multi-Offset Detector-Native ReID Readout for Efficient UAV Multi-Object Tracking

Defeng Sun, Shaowu Dai, Zhicai Xiao, Shun Sun

Joint detection-and-embedding (JDE) trackers avoid per-detection crop inference by reading identity features for re-identification (ReID) from the detector. The detection-center readout, however, does not use a track prediction when forming the appearance descriptor. This is restrictive in unmanned aerial vehicle (UAV) video, where small targets, camera motion, and localization jitter can displace the detection center from stable identity features. We propose the Motion-Guided Multi-Offset Detector-Native ReID Readout. It samples shared feature pyramid network (FPN) identity maps at the current detection and an assigned reference box obtained from globally compensated track predictions. A 12-D detection-level motion prior conditions the Motion-Guided Dense Branch. The Multi-Offset Token Branch retains source, pyramid-level, and box-relative offset information before the two branches produce one normalized appearance descriptor. Controlled ablation on the VisDrone2019-MOT Val split shows that the full dual-branch readout improves identity association compared with the detection-center readout. On VisDrone2019-MOT test-dev, the complete system achieves an Identification F1 score (IDF1) of 67.16 with HybridSORT. Its Multiple Object Tracking Accuracy (MOTA) is 53.15, with a throughput of 16.98 frames per second (FPS). With Deep OC-SORT, it achieves 66.32 IDF1, 52.19 MOTA, and 17.07 FPS. A candidate-conditioned pair residual provides a HybridSORT application extension, further raising IDF1 to 67.53.