DOI: 10.3390/s26196180 ISSN: 1424-8220

Quantifying Subject-Identity Variance in Spectral EEG Features and Its Role in Machine Learning Evaluation Leakage

Hassan Ugail, Richard Wirt, Newton Howard

Electroencephalography (EEG) is widely used to study cognitive states and clinical biomarkers, yet the contribution of stable between-subject differences to common EEG feature representations is rarely quantified directly or incorporated into evaluation design. Here, we examine subject-linked structure in spectral EEG features using a longitudinal four-session dataset spanning 199 days and replicate key findings across four public datasets covering 39 to 395 participants, multiple paradigms, and recording systems ranging from low-channel consumer devices to research-grade EEG. In spectral band-power representations, subject identity accounted for substantially more variance than session-related drift and supported extremely strong individual discriminability under subject-wise held-out protocols, with high verification performance and robust cross-session recognition over months. However, this identity structure was not fully invariant across paradigms, showing stronger transfer in richer, higher-channel recordings than in low-channel consumer EEG under larger task shifts. We further show that the same subject-linked structure can inflate downstream clinical classification when evaluation is performed with naive segment-level splits, demonstrating that apparent diagnostic performance can partly reflect identity leakage rather than biomarker learning. These findings highlight subject identity as a measurable and persistent source of structured variance in spectral EEG features and support routine use of subject-wise evaluation and explicit confound diagnostics in EEG machine-learning studies.