MSA-CNN: A Multi-Scale Attention Convolutional Neural Network for fNIRS-Based Emotion Recognition
Deping Huang, Xiu Zhang, Ye Li, Jingfu Wu, Youzhi YueFunctional near-infrared spectroscopy (fNIRS) has attracted increasing attention in affective brain–computer interface research due to its non-invasive nature, portability, and robustness to motion artifacts. However, substantial inter-subject variability in neural responses remains a major challenge for subject-independent emotion recognition. To address this issue, this work presents an effective integration of multi-scale temporal convolution and dual-attention mechanisms for subject-independent fNIRS emotion recognition evaluated under the leave-one-subject-out protocol within a single dataset. The proposed framework employs multi-scale temporal convolutions to capture hemodynamic characteristics at different temporal resolutions and incorporates channel and temporal attention mechanisms to adaptively emphasize informative brain regions and critical temporal segments. Experiments were conducted on both a self-collected fNIRS emotion dataset and the publicly available ENTER dataset using the Leave-One-Subject-Out (LOSO) evaluation protocol. On the self-collected dataset, MSA-CNN achieved an accuracy of 65.06 ± 7.10% with an F1-score of 0.605. On the ENTER dataset, the proposed model obtained an accuracy of 68.91% and an F1-score of 0.621, outperforming conventional machine learning approaches and several representative deep learning baselines. Ablation studies further demonstrated the positive contributions of both the multi-scale convolutional structure and the dual-attention mechanism. Experimental results on both the self-collected and ENTER datasets demonstrate that the proposed MSA-CNN achieves competitive emotion recognition performance under the LOSO protocol. Class-wise evaluation using precision, recall, and the F1-score further provides a comprehensive assessment of the model’s classification behavior. These results indicate the effectiveness of the proposed framework for cross-subject fNIRS-based emotion recognition under the current experimental settings. The results indicate that multi-scale temporal feature learning combined with attention mechanisms can effectively enhance fNIRS-based emotion recognition performance and provides a promising framework for within-dataset cross-subject evaluation in fNIRS-based emotion recognition. Future work will focus on expanding the subject population, conducting cross-dataset train–test evaluations, and incorporating multimodal neural signals to further improve robustness and generalization.