DOI: 10.3390/electronics15194483 ISSN: 2079-9292

Aligning Acoustic and Semantic Representations for Speech Aspect-Based Sentiment Analysis

Yifei Miao, Zhongqing Wang, Guodong Zhou

Fine-grained sentiment information about specific aspects is the target of aspect-based sentiment analysis (ABSA). Nevertheless, when real-world situations involve sentiment expressed through multiple modalities, traditional text-centric approaches are prone to information loss and ambiguity. In this work, we introduce a novel speech aspect-based sentiment analysis task. By taking raw speech signals as input, this task outputs aspect terms, opinion expressions and sentiment polarities associated with each aspect and thereby leverages the rich information in speech, including paralinguistic cues such as prosody, stress and intonation, to strengthen sentiment analysis. To facilitate controlled evaluation, we construct dedicated speech ABSA datasets from two domains (Restaurant and Laptop) with roughly 6 h of audio and over 9000 annotated quadruples. Text-to-speech synthesis is employed for training and validation sets in our data construction strategy, while human-recorded speech is reserved for test sets to assess generalisation under read-speech conditions. Furthermore, we adopt a cross-modal mixup approach based on optimal transport to jointly extract sentiment elements from both speech and ASR-generated text and effectively align acoustic and semantic representations. Experimental results on both domains confirm the significance of the suggested speech aspect-based sentiment analysis task, with our model reaching F1 scores of 45.28% and 40.06% on Restaurant and Laptop, respectively, and outperforming several strong audio-only, ASR-based and multimodal baselines.