Evaluating Validation Strategies in Motor Imagery EEG: A Full-Cohort GAF–PLV Analysis and Matched Sensitivity Study
Wenwen Chang, Hesam Akbari, Muhammad Tariq Sadiq, Renjie Lv, Rab NawazBackground: Performance estimates in motor-imagery electroencephalography (MI-EEG) can depend strongly on how observations are partitioned for training, model selection, and testing. Random sample- or window-level splitting may place data from the same participant in different folds and therefore does not answer the same question as evaluation on previously unseen participants. Methods: We evaluated a previously developed Gramian angular field–phase-locking value (GAF–PLV) classifier on the retained full cohort (N=105) using binary left-versus-right MI and leave-one-subject-out cross-validation (LOSO). Separately, a predefined, outcome-independent subset (N=30) was used for a matched sensitivity analysis of eight classifiers, binary and four-class tasks, and three validation strategies: random five-fold cross-validation, LOSO, and nested LOSO with subject-grouped inner model selection. Results: In the full-cohort GAF–PLV analysis, mean accuracy was 58.07% ± 8.27% and Macro-F1 was 53.48% ± 11.19%, with substantial between-subject variability. In the predefined matched subset, performance estimates and numerical model rankings changed across validation strategies. For example, the numerically highest binary-accuracy model was ShallowConvNet under random five-fold cross-validation, DeepConvNet under LOSO, and ATCNet under nested LOSO. Conclusions: Random within-cohort classification and generalisation to previously unseen subjects are distinct evaluation targets. MI-EEG reports should state the cohort, partition unit, validation design, and model-selection procedure. Rankings in the multi-model analysis are conditional on the predefined 30-subject subset and common 0.4 s input setting and are not presented as definitive full-cohort or architecture-optimal rankings.