A Stability and Accuracy Evaluation of CNN, LR and GA-BP Models for Pepper Leaf Disease Recognition Based on a Multi-Dimensional Visual Feature Dataset
Xueting Ma, Yifei Li, Na Jia, Xiaodong Xu, Fuxiang Lei, Ganggang Guo, Kaijie QiPepper suffers from bacterial leaf spot and yellow leaf curl, which substantially reduce crop yield and fruit quality. Traditional manual diagnosis suffers from delayed response, subjective bias and heavy labor consumption, while existing intelligent detection pipelines lack standardized preprocessing workflows, quantitative feature screening and systematic model comparison. To fill these research gaps, we built a pepper leaf dataset with 1260 samples (healthy, bacterial spot, yellow leaf curl). Three segmentation algorithms (Lab b-channel, RGB super-green, Otsu-ACWE) were quantitatively assessed to select the optimal preprocessing scheme. We extracted 32 fused visual features (27 RGB/HSV/Lab color moments + five gray-level co-occurrence matrix (GLCM) texture metrics) and adopted a random-forest classifier to eliminate seven low-contribution redundant features, retaining 25 discriminative variables. Three representative models, namely convolutional neural network (CNN), logistic regression (LR), and genetic-algorithm-optimized back-propagation neural network (GA-BP), were constructed for parallel comparison via 20 independent repeated trials, with accuracy, precision, recall, F1-score and area under the receiver operating characteristic curve (AUC) as evaluation indicators. The results verified that Lab b-channel segmentation achieved superior background separation and intact lesion edge retention. CNN yielded the best performance, with an average test accuracy of 97.67% and an average AUC of 0.999, accompanied by minimal metric standard deviations and outstanding stability. LR exhibits low computational cost and fast training, which is promising for applications with limited computing resources. In contrast, GA-BP shows weak nonlinear fitting ability and severe prediction fluctuations, making it unsuitable for high-precision diagnosis. This study proposes a standardized experimental framework to offer theoretical guidance and algorithmic references for intelligent vegetable leaf disease identification. All experiments were conducted on a dataset collected under standardized indoor single-illumination conditions; therefore, the conclusions of this study are only applicable to such controlled scenarios.