DOI: 10.3390/info17100955 ISSN: 2078-2489

Hierarchical Transformer for Classifying Depressive-Episode Conditions from Wrist Actigraphy: A Preliminary Participant-Level Study

Braj Kishor Pathak, Abhinav Shukla, Ayush Kumar Agrawal, R Kanesaraj Ramasamy, Parul Dubey

Depression is increasingly investigated through wearable digital phenotyping because alterations in motor activity, daily routines, and circadian organisation may provide objective behavioural information relevant to mental-health screening. Wrist actigraphy enables continuous, non-invasive monitoring of such longitudinal activity patterns in natural settings. However, existing machine- and deep-learning approaches often rely on engineered summary features or isolated temporal windows, limiting representation of within-day and between-day behavioural variation. Small participant-level datasets also increase the risk of information leakage and unreliable probability estimates. This retrospective secondary-data proof-of-concept study analysed minute-level wrist actigraphy from 55 participants, comprising 23 participants experiencing unipolar or bipolar depressive episodes and 32 healthy controls. Model development and internal validation used nested leave-one-subject-out cross-validation, with preprocessing, self-supervised representation learning, hyperparameter selection, and calibration restricted to training participants within each fold. The novelty lies in integrating hierarchical temporal modelling, leakage-resistant representation learning, interpretable circadian information, and uncertainty-aware prediction within a single activity-only framework. Performance was evaluated using balanced accuracy, sensitivity, specificity, F1-score, MCC, AUROC, AUPRC, and calibration measures. The five-seed probability-averaged HC-MAT ensemble achieved balanced accuracy of 0.872, sensitivity of 0.870, specificity of 0.875, F1-score of 0.851, MCC of 0.741, and AUROC of 0.920; across individual seeds, balanced accuracy was 0.859 ± 0.010. Although HC-MAT achieved the highest numerical performance across several metrics, none of its comparisons with baseline or ablated models remained statistically significant after Holm correction. These results support HC-MAT as a proof-of-concept framework requiring validation in substantially larger, independent clinical cohorts before clinical deployment.