DOI: 10.3390/app16189290 ISSN: 2076-3417

Speech State Analysis Based on Self-Supervised Representation Shift Under Complex Speech Interference

Furui Zhao, Zhihui Xu, Xingwei Zhang, Hengrui Guo, Zhihu Wang, Jianwei Niu

Safety-critical control environments, such as nuclear power plant main control rooms, involve multi-operator collaboration, continuous information exchange, and stringent reliability requirements, making operator-state monitoring important for task performance and system safety. Speech is naturally produced, non-intrusive, and readily available for continuous acquisition, but multi-operator settings introduce substantial speech interference, while high-arousal or abnormal-state speech is difficult to collect at scale. This study proposes a speech-state analysis method based on self-supervised representation shift. Low- and high-arousal speech conditions are constructed from a public emotional speech dataset using the valence–arousal framework. Emotion2Vec extracts high-level embeddings, an individual low-arousal baseline is established, and cosine distance quantifies representation shift. The method is compared with a convolutional neural network (CNN) and a CNN with long short-term memory (CNN+LSTM) under clean speech, white-noise interference, and multi-speaker speech interference; supervised-model dependence on training set size is also examined. Low- and high-arousal speech remain partially separable in representation space, while the proposed method shows less performance degradation than the supervised baselines under interference, particularly multi-speaker interference. These findings support individual baseline representation shift as a feasible approach for non-intrusive state monitoring in safety-critical multi-operator environments.