DOI: 10.3390/agriengineering8090402 ISSN: 2624-7402

An Image-Based Approach for Grape Yield Estimation During Mechanical Harvesting Using Pre-Trained Visual Embeddings

Giuliano Vitali, Federico Terzi, Marco Arru, Francesco Pelusi, Cristiano Fragassa, Matteo Gatti, Michele Mattetti

Accurate yield mapping is essential for precision viticulture, but direct yield measurement during mechanical harvesting remains difficult because of the heterogeneous harvested biomass. This study evaluates the use of images acquired during harvesting to estimate grape yield through features extracted from pre-trained deep learning models. Images were collected over two seasons in a commercial Merlot vineyard using action cameras mounted on a grape harvester. Five pre-trained models, including one convolutional neural network and four Vision Transformer architectures, were compared. The extracted features were related to measured yield using regularized regression models. The results show that pre-trained visual features can reliably estimate grape yield despite the presence of berries, grape clusters, leaves, shoots and motion blur. The convolutional neural network achieved prediction accuracy comparable to the Vision Transformer models, indicating that local image features remain highly informative for biomass-flow estimation. A sparse fusion approach identified more stable features, supporting improved model stability under the investigated conditions. This work demonstrates that general-purpose pre-trained visual embeddings provide an effective basis for grape yield estimation without task-specific model training. The proposed method, not depending on dedicated sensors, specific calibration or model tuning, offers a practical and scalable solution for integration into commercial grape harvesters, including aftermarket installations.