Multichannel Polysomnographic Sensor Fusion for Automatic Sleep Stage Classification
Simone Mari, Lorenzo Di Vaira, Giovanni Bucci, Fabrizio Ciancetta, Edoardo FiorucciAutomatic sleep stage classification can reduce the burden of manual polysomnographic scoring, but reliable models must account for heterogeneous physiological signals, temporal dependencies, and severe class imbalance. This study presents a dual-stream deep learning framework for five-class sleep staging from 30 s polysomnographic segments using two EEG signals, one EOG signal, and one submental EMG signal. Time-domain features are extracted through one-dimensional convolutions, while complementary time–frequency information is obtained from short-time Fourier transform spectrograms processed by a two-dimensional convolutional branch. The resulting features are fused and analyzed over a five-segment context window using a bidirectional long short-term memory network. Prolonged wake periods are censored only during training to reduce majority-class dominance without altering the evaluation distribution. Subject-wise five-fold cross-validation across 78 subjects yielded a mean accuracy of 91.5%, a balanced accuracy of 78.5%, a macro-F1 score of 77.4%, and a Cohen’s kappa of 0.829. Controlled architectural and channel-ablation experiments further quantified the contributions of temporal context, time–frequency processing, wake censoring, and the individual PSG signals. A lightweight time-domain-only variant achieved 88.8% accuracy, 72.7% balanced accuracy, a macro-F1 score of 73.8%, and a Cohen’s kappa of 0.779. These findings support the use of multimodal time–frequency fusion for robust automatic sleep staging while highlighting the trade-off between classification performance and model complexity.