DOI: 10.3390/agronomy16151478 ISSN: 2073-4395

Self-Supervised Multimodal Learning for Preharvest and Postharvest Fruit Quality Assessment Using Images and Environmental Sensors

Chuhuang Zhou, Tanghua Wang, Xin Zeng, Fei Wang, Fanfei Meng, Zheng Yang, Min Dong

Fruit preharvest–postharvest quality assessment is essential for precision harvesting, intelligent grading, storage management, and supply-chain loss reduction. However, conventional approaches mainly rely on single-point postharvest inspection and cannot adequately capture the long-term effects of preharvest fruit phenotypes and environmental dynamics on quality formation. To address limited prediction accuracy under few-label conditions, insufficient multimodal fusion, and weak cross-orchard generalization, this study proposes FruitSSL-QNet, a self-supervised multimodal learning framework for jointly modeling preharvest fruit images, environmental sensor time series, and postharvest quality indicators. The framework employs visual masked reconstruction to learn fine-grained phenotype features, including color, texture, lenticel distribution, disease spots, and maturity patterns. Environmental temporal masked modeling is used to capture the cumulative effects of temperature, humidity, light intensity, soil moisture, and rainfall. Bidirectional cross-attention, gated fusion, and contrastive alignment are further integrated to learn complementary and semantically consistent image–environment representations. Experimental results demonstrate that FruitSSL-QNet outperforms SVM, Random Forest, XGBoost, LSTM, GRU, TCN, Transformer, and MM-Transformer across multiple quality assessment tasks. The proposed model achieves a maturity recognition accuracy of 89.6%, exceeding MM-Transformer by 4.4 percentage points. Compared with the corresponding baseline results, the prediction errors for sugar content, firmness, and shelf life are reduced by 21.1%, 22.2%, and 22.5%, respectively. The decay-risk AUC reaches 0.921, and the cross-site F1-score reaches 0.867, indicating strong risk discrimination and stable generalization across orchard environments. Ablation experiments further confirm the contributions of visual self-supervision, environmental temporal self-supervision, and cross-modal alignment.

More from our Archive