DOI: 10.3390/diagnostics16162609 ISSN: 2075-4418

Automated Multimodal Sleep Staging Using DWT-Based Wavelet Decomposition and Explainable Machine Learning with Signal Sculpting Topographies

Adnan Sami Sarker, Kazi Mahatir Mohammed Samir, Zunayed Khan Shakib, Md Kishor Morol, Tze Hui Liew

Objectives: Sleep staging from polysomnographic (PSG) recordings is clinically critical for diagnosing sleep-related disorders, yet manual scoring by certified technologists remains time-consuming, costly, and subject to inter-rater variability. Methods: This study presents an automated, explainable, and multimodal framework for five-class sleep stage classification using simultaneously acquired electroencephalography (EEG), electrooculography (EOG), and electromyography (EMG) signals. A total of 1946 annotated 30 s epochs from 30 healthy adult recording sessions (Sleep-EDF Expanded and Sleep Cassette subset) were processed through a 37-dimensional multimodal feature extraction pipeline encompassing temporal amplitude statistics, frequency-domain spectral band powers, nonlinear entropy and complexity measures, and Daubechies-4 discrete wavelet transform (DWT) energy coefficients. Four classical machine learning classifiers -Random Forest (RF), Support Vector Machine with radial basis function kernel (SVM-RBF), Gradient Boosting (GB), and K-Nearest Neighbours (KNN, k = 7) were benchmarked under stratified five-fold cross-validation. Results: SVM-RBF achieved the highest macro-averaged F1-score of 0.7322 (Cohen’s kappa 0.6784, overall accuracy 75.18%). N3 deep slow-wave sleep achieved the highest per-class F1 of 0.879, while N1 light sleep was the most challenging (F1 = 0.668). SHapley Additive exPlanations (SHAP) and RF mean decrease in Gini impurity (MDGI) analysis jointly identified EMG root mean square amplitude (MDGI = 0.0805), gamma band power (0.0784), and permutation entropy (0.0434) as the three most discriminative features. As a novel methodological contribution, sixteen categories of signal sculpting visualisations were developed, translating abstract multivariate features into clinically interpretable graphical representations. Conclusions: The proposed framework achieves substantial kappa agreement approaching the lower bound of expert inter-rater reliability (0.76–0.82) while providing full model transparency, with direct implications for wearable sleep monitoring device design.

More from our Archive