Synthetic Data for Data-Efficient Agricultural Computer Vision: Three Datasets and a Multi-Task Benchmark Across Real and Simulated Domains
Steven Moonen, Wouter Jansen, Wenzhi Liao, Mohammad Hasan Rahmani, Ayyoub Ahar, Abdellatif Bey-Temsamani, Jan Steckel, Nick MichielsThe availability of large annotated datasets remains a major bottleneck for deploying computer vision systems in agricultural and agri-food applications, particularly for tasks requiring fine-grained annotations such as defect detection and instance segmentation. Synthetic data generation offers a scalable alternative, but its effectiveness in complex real-world scenarios remains insufficiently understood. In this work, the role of synthetic data is investigated across three representative tasks each posing unique challenges: apple detection in orchard environments, potato–stone classification in industrial sorting, and carrot crack detection for quality inspection. The real and synthetic data, as well as their annotations, are publicly available. The synthetic data is created using a set of tools for fast large-scale procedural scene generation with highly detailed natural assets. For each task, datasets are constructed combining real and synthetically generated images with pixel-level annotations. Training strategies are systematically evaluated using fully real, limited real, synthetic-only, and combined datasets. The results show that synthetic-only training leads to a clear performance gap on real-world data, with mAP50 decreasing from 0.866–0.891 for fully real training to 0.396–0.703 for synthetic-only training across the three use cases. However, combining synthetic data with only 10 real images recovers 75–98% of the gap between training on those 10 real images alone (mAP50 0.194–0.497) and full real training, reaching mAP50 values of 0.723–0.885. These findings highlight the potential of synthetic data for improving data efficiency and provide a reproducible benchmark for future research on sim-to-real transfer in agricultural computer vision.