DOI: 10.3390/s26155001 ISSN: 1424-8220

Bus-Mounted Vision Sensing for Traffic Object Detection: BFTD and a Local–Global Attention Framework

Wenjing Gao, Nan Zou

Bus-mounted vision sensing provides a practical and complementary perspective for intelligent transportation systems, but reliable traffic object detection from bus front-view cameras remains challenging because elevated viewpoints induce severe scale skewness, dense interactions around bus stops and intersections, and frequent heterogeneous occlusion. To support this sensing scenario while avoiding ambiguity with previously used dataset acronyms, we construct the Bus Front-view Traffic Dataset (BFTD), a high-resolution benchmark collected from forward-facing cameras mounted on multiple buses operating on urban routes during real-world service. The BFTD contains 8131 images and 56,137 annotated instances across five traffic-participant categories, covering dense pedestrians, mixed-traffic flow, illumination variation, rain, fog, and occlusion-prone scenes. Based on the visual characteristics of bus-mounted cameras, we propose YOLO-M2LA, a local–global attention detection framework in which CBS-SPD preserves fine-grained information during early downsampling and M2LA couples multi-scale local context modeling with efficient global dependency aggregation. Extensive experiments on BFTD and public benchmarks show that the proposed framework improves detection accuracy, particularly for small and visually crowded traffic participants, while maintaining a practical accuracy–efficiency trade-off. Dataset statistics, condition-specific evaluation, ablation analysis, and qualitative visualization further support the effectiveness of BFTD and YOLO-M2LA for vision-based traffic sensing. The dataset and implementation are publicly available online.

More from our Archive