DOI: 10.3390/smartcities9100160 ISSN: 2624-6511

Benchmarking Recent YOLO Architectures for Detecting Vehicles, Pedestrians, and Motorcycles at Complex Urban Intersections: Evidence from a Hybrid Mexico City–LISA Dataset

Julio Saucedo-Soto, Viridiana Hernández-Herrera, Moisés Márquez-Olivera, Antonio-Gustavo Juárez-Gracia, Octavio Sánchez-García, Amadeo Argüelles-Cruz

Urban road user detection supports safety-critical applications in Intelligent Transportation Systems, but architecture assessment requires balancing predictive performance with computational efficiency under controlled and leakage-aware evaluation conditions. This study adopts a three-stage benchmarking framework for vehicle, motorcycle, and pedestrian detection using a hybrid dataset of 4983 images derived from LISA and traffic scenes captured at 18 complex intersections in Mexico City. Stage I provides an exploratory comparison of six medium-sized YOLO architectures (YOLOv8m, YOLOv9m, YOLOv10m, YOLOv11m, YOLOv12m, and YOLOv26m), whereas Stage II restricts leakage-controlled confirmatory evaluation to three selected architectures: YOLOv10m, YOLOv11m, and YOLOv26m. Stage III examines the corresponding YOLOv10n and YOLOv26n nano variants. Both sources were manually annotated under a unified three-class protocol, producing 3483 quantitative images with 6624 ground truth instances and 1500 images for qualitative inference. Annotation reproducibility reached a mean pairwise IoU of 0.85 ± 0.11 and Fleiss’ κ of 0.968. A retrospective audit of the original image-level split identified acquisition group overlap and confirmed near duplicates across partitions. A corrected group-disjoint split eliminated shared groups, exact duplicates, and confirmed near duplicates while preserving the aggregate source and class distributions. Under three-seed leakage-controlled confirmatory evaluation, YOLOv26m achieved the highest mAP@0.50:0.95 among the three Stage II confirmatory models (0.5717 ± 0.0013), followed by YOLOv11m (0.5632 ± 0.0012) and YOLOv10m (0.5516 ± 0.0014), with paired bootstrap analysis supporting all three pairwise differences after Holm correction. Source-specific evaluation revealed consistently lower performance on the Mexico City subset than on LISA. YOLOv11m showed the lowest overall calibration error (ECE = 0.052 ± 0.003). At the nano-scale, YOLOv10n provided higher mAP@0.50:0.95, lower latency, and greater throughput than YOLOv26n, whereas YOLOv26n required fewer GFLOPs and less benchmark energy. The Stage II group-disjoint protocol produced more conservative performance estimates than Stage I; however, the observed mAP reduction cannot be quantified as the effect of leakage removal alone because the two stages also differed in optimization configuration and repeated-run design. These findings support multi-objective model selection based on predictive performance, calibration, latency, complexity, memory, energy consumption, and deployment constraints.