DOI: 10.3390/automation7040130 ISSN: 2673-4052

Vision-Map Fusion Multi-Object Tracking at Complex Intersections Using HD Map Priors and Nonlinear Filtering

Dezheng Ma, Lan Tang

Accurate multi-object tracking and metric localization support traffic monitoring and cooperative intelligent transportation at complex intersections. This study presents a fixed-camera vision-map fusion framework that addresses two practical difficulties: axis-aligned boxes poorly represent turning vehicles, and unconstrained image-plane tracking can produce physically implausible trajectories. A map-aided frontend first generates candidate detections using improved You Only Look Once version 8 nano (YOLOv8n) horizontal bounding box (HBB) branch and an improved YOLOv8 oriented bounding box (OBB) branch. A high-definition (HD) map selector then retains the candidate geometry consistent with the straight-driving or turning region and converts it into a unified detection record. The selected reference point is projected to the ground plane through an offline-estimated homography, whereas the appearance feature bypasses the homography and is passed directly to the association stage. The tracking backend uses a 12-dimensional joint image/metric state, symmetric central-difference evaluations of the process and measurement functions, appearance-motion association, and a feasible-road projection derived from HD-map lane polygons. On the evaluated public sequences, the complete configuration achieved a multiple object tracking accuracy (MOTA) of 74.5%, an identification F1 score (IDF1) of 82.6%, 614 identity switches, and a throughput of 26.8 frames per second (FPS) on an RTX 4090 workstation. In a descriptive Vehicle-in-the-Loop case study involving one instrumented vehicle at one intersection, the overall localization mean absolute error (MAE) was 0.180 m, compared with 0.208 m for the baseline end-to-end configuration. These results indicate the feasibility of combining branch-specific vehicle geometry with map-constrained tracking; controlled same-detector comparisons, repeated multi-vehicle trials, and embedded-device latency and power profiling remain necessary for broader claims.

More from our Archive