DOI: 10.1049/ipr2.70453 ISSN: 1751-9659

DPCLT: Dynamic Position‐Calibrated Lightweight Transformer for Efficient and Robust Visual Object Tracking

Rong Wang, Zhenhuai Gao, Xiaoling Gao

ABSTRACT

Transformer‐based visual tracking has achieved strong representation capability, yet balancing tracking accuracy, temporal robustness and real‐time efficiency remains difficult under limited computational resources. The challenge is particularly pronounced in dynamic scenes, where lightweight feature extraction can weaken spatial discrimination, while inaccurate cross‐frame associations may accumulate into tracking drift. To address this problem, this paper proposes a dynamic position‐calibrated lightweight Transformer (DPCLT) that organizes efficient feature modelling, temporal representation calibration and adaptive inference within a unified tracking process. The framework first models spatial structure and channel semantics through a compact dual‐correlation mechanism, preserving discriminative information without relying on computationally expensive global attention. It then establishes a position–memory calibration process in which foreground and background prototypes are used to refine historical representations, and the calibrated temporal information is further converted into spatial guidance for current‐frame localization. This bidirectional interaction improves cross‐frame correspondence and reduces the influence of unreliable memory updates. In addition, scene complexity and target motion are incorporated into inference‐path selection, enabling the tracker to adjust its computational cost across frames. Experiments on LaSOT, GOT‐10k and TrackingNet show that DPCLT achieves a competitive accuracy–efficiency trade‐off and improves tracking stability in challenging conditions including occlusion, fast motion and background interference.

More from our Archive