Saturated In-Domain, Separable Out-of-Domain: A Hierarchical Multi-Scale Lesion-Attention Network and a Zero-Shot Cross-Corpus Protocol for Multi-Crop Leaf Disease Classification
Songul Karakus, Mehmet Burukanli, Davut AriLeaf disease classifiers are ranked by held-out accuracy on their training corpus, where that number now exceeds 98% for many modern backbones. We ask what such a measurement can still resolve. On a four-crop, 21-class corpus of 7179 de-duplicated images, we train 26 architectures under one protocol with three seeds and shared partitions, and we execute the entire benchmark twice. The 26 architectures span 1.68 percentage points (pp) of internal macro-F1. The same configuration differs by 0.34 pp on average and by up to 1.06 pp between the two executions, and the internal ordering only partly replicates (Spearman’s ρ=0.69, with VGG-19 moving from 17th to 1st). Applied zero-shot to an independently collected mango corpus, the same checkpoints spread over 47 pp of accuracy and their ordering does replicate (ρ=0.92). We propose HMLA-Net, which adds multi-scale fusion, gated multi-head lesion-aware pooling, an auxiliary crop head and MixStyle to a CAFormer-S18 backbone and averages the posterior over the Klein four-group of flips at inference. In the reference execution, it attains the highest internal macro-F1 (0.9845) and the highest zero-shot accuracy (0.699) of the 26 architectures, with the smallest degradation (28.9 pp) and the lowest rate of predictions leaking into labels absent from the target corpus. The second execution reproduces its zero-shot rank but not its internal one. Five backbones retrained with its complete recipe show that its margin over its own backbone lies within replicate variation, and in a 16-removal ablation measured on both partitions, only the large effect of ImageNet pretraining remains consistent across the two executions. A cross-corpus duplicate audit finds no exact duplicates and no evidence of substantial overlap, and one class, mango sooty mould, collapses to near-zero recall in 19 of 26 architectures. The contribution is the protocol: replicate the benchmark, evaluate out of domain, and report results per class.