A Comparative Evaluation of Deep Learning Architectures for Weed Classification, with an Exploratory Out-of-Distribution Analysis of Albanian Field Images
Rajesh Kumar, Narasimha Rao Vajjhala, Ervin Ramollari, Earta JocaAutomated weed classification supports site-specific weed management by enabling targeted interventions and reducing environmental and operational costs. This study presents a multi-seed evaluation of six deep learning configurations—YOLO26n-cls, Vision Transformer (ViT-B/16), DINOv2 (ViT-S/14) with linear probing and full fine-tuning, ResNet-50, and EfficientNet-B0—for nine-class weed classification using a stratified subset of the DeepWeeds benchmark. The models were evaluated using a leakage-checked train/validation/test split and repeated training across multiple random seeds to assess performance stability and statistical reliability. Among the evaluated approaches, DINOv2-FT and EfficientNet-B0 demonstrated the strongest overall performance, while statistical analysis showed that differences among several high-performing models were not consistently significant across seeds. The study also identifies the importance of appropriate transfer-learning strategies and hyperparameter selection, demonstrating that model performance can be strongly affected by optimisation choices. A reproducibility issue affecting ViT-B/16 was identified and transparently reported, leading to its exclusion from multi-seed statistical comparisons. An exploratory out-of-distribution probe using unlabelled Albanian field images further highlights the challenges of geographic domain shift and the need for locally collected, labelled datasets before reliable regional deployment. Overall, this work provides a systematic comparison of modern deep learning architectures for weed classification and emphasises reproducibility, statistical validation, and careful interpretation of model rankings.