Temporal and Feature Alignment in V2I Fusion for 3D Object Detection and Tracking at Intersections
Jingxiong Meng, Junfeng ZhaoABSTRACT
Vehicle‐to‐infrastructure (V2I) cooperative perception is important for safer intersection operation because roadside sensors can reveal road users that are occluded, distant, or poorly observable from onboard cameras alone. Dependable V2I fusion is hindered by two practical issues. First, sensing, processing, and wireless transmission delays introduce temporal misalignment between infrastructure and vehicle streams. Second, large viewpoint differences create feature‐level domain discrepancy that degrades cross‐view matching and fusion. This paper proposes a unified alignment framework with a feature alignment module (FAM) and a temporal alignment module (TAM) to improve cross‐view consistency before fusion. FAM reduces the discrepancy between infrastructure and vehicle object queries to improve matching reliability, while TAM compensates delayed infrastructure observations through lightweight motion‐based temporal alignment. Experiments on V2X‐Sim and DAIR‐V2X demonstrate consistent improvements over representative baselines under both ideal and delayed settings. Compared with the vehicle‐only DQTrack baseline, it improves bicycle and motorcycle AP by 3.1 and 4.8 points. On the real‐world DAIR‐V2X benchmark, it further improves cyclist detection AP and tracking AMOTA by 2.7 and 3.4 points. These results indicate that aligning heterogeneous and asynchronous V2I information before fusion can improve both benchmark accuracy and the practical reliability of cooperative perception in safety‐critical intelligent transportation scenarios.