DOI: 10.3390/bdcc10100328 ISSN: 2504-2289

Cross-Regional Benchmarking of Enterprise Cloud Object Recognition Services: A Comparative Analysis of AWS, Azure, Google Cloud Vision, and YOLO12

Emmanuel Masa-Ibi, Kofi Nyarko

This paper benchmarks AWS Rekognition, Azure Computer Vision, Google Cloud Vision, and YOLO12 on a pre-specified subset of GeoDE, comprising 12 object categories, 16 countries, and six regions. Applying the same country exclusion to the completed export yields 17,676 images by aggregate count for every system. Image-weighted exact-plus-related recognition accuracies are 94.93% for AWS, 84.25% for Azure, 85.03% for Google, and 58.32% for YOLO12. Exact-only scoring reduces these values to 83.89%, 76.79%, 72.74%, and 58.01%, respectively, reversing the Azure–Google ordering and demonstrating sensitivity to semantic matching. On the five stated vocabulary-overlapping categories (7524 images), YOLO12 achieves 95.28% compared with AWS at 99.22%. Reported local CPU inference time is lower than cloud API round-trip time, but the timing boundaries differ. European images have the highest descriptive accuracy for the cloud services; YOLO12 is highest on Southeast Asian images. The aggregate export does not establish image-level pairing or bootstrap units, so unverifiable original inferential results are withdrawn rather than treated as image-level evidence. No power analysis establishes the cause of the earlier nonsignificant regional results. Lower pooled recognition of stall and lighter does not establish a cultural or geographic mechanism. Proprietary training data were not analyzed. The separately reported human baseline (500 images) achieved 97.2% accuracy with Fleiss’ κ=0.94. Results combine November 2025 records with 10 September 2026 completion runs and should inform deployment-specific validation rather than universal provider rankings.