Leakage-Free Benchmarking of Electronic Noses for Beef Freshness: A Signal-Richness Criterion for Model Selection
Erkan Caner OzkatLow-cost metal-oxide-semiconductor (MOS) electronic noses promise rapid, non-destructive meat freshness screening, and published classifiers frequently approach perfect accuracy. Such figures are rarely tested against the two conditions that most inflate them: a target-derived label among the inputs, and random splitting of the correlated samples. Beef freshness is benchmarked here on a public 11-sensor, 12-cut MOS dataset using leakage-free leave-one-cut-out cross-validation in order to predict freshness class and total viable count (TVC) with paired significance tests. A gradient-boosted-tree pipeline is the strongest model (accuracy 0.81±0.10, macro-F1 0.68±0.15, TVC R2=0.77), significantly outperforming a multi-scale attention convolutional network (macro-F1 0.50±0.15; p<0.001). The advantage of this study lies in the representation, not the model family: a network given the same window summaries reaches 0.64±0.17, indistinguishable from the tree. Near-perfect accuracy returns only when TVC is supplied as a feature or samples are split at random (macro-F1 0.97). Under nested, per-fold selection, a five-sensor subset matches the full array. On a rich BME688 heater profile dataset, the network surpasses the tree, an advantage that vanishes as the profile shortens to one step. Evaluation and representation, not architecture, govern reported performance; a signal-richness criterion predicts when a deep temporal model is justified.