DOI: 10.3390/bdcc10080276 ISSN: 2504-2289

Accurate and Robust Multimodal Emotion Recognition for Human–Robot Interaction via Dynamic Graph Learning with Pairwise Cross-Modal Alignment

Xinyang Zhou, Jiahao Wu, Hongming Xu, Jinghan Mei, Zeyang Chen, Junxiong Zhang, Yitong Chen, Yanrui Jin, Chengliang Liu, Chenggang Yuan

Multimodal emotion recognition in conversation (MERC) aims to identify the emotions in each utterance by modeling textual, acoustic, and visual evidence. Compared with unimodal emotion recognition in conversation (ERC), MERC can leverage complementary textual, acoustic, and visual information to support more accurate and consistent emotion inference. However, coordinating intramodal contextual dependencies, cross-modal alignment, and temporal affective dynamics within a unified framework in MERC is challenging. Existing solutions have advanced MERC through contextual modeling, multimodal fusion, and graph-based reasoning, but they still often rely on static relational assumptions or stage-wise coordination of modalities. This limits their ability to jointly model fine-grained relations, selective cross-modal interactions, and dynamic changes in emotion. To address these issues, we propose DGL-PCA (dynamic graph learning with pairwise cross-modal alignment), a dynamic graph-based framework for MERC. The model combines time-aware relation construction, dynamic time-aware heterogeneous graph modeling, and pairwise cross-modal alignment to improve prediction accuracy. This coordinates temporal affective dynamics, structured dialogue context, and multimodal interaction more explicitly than conventional coarse fusion or static graph formulations. Extensive experiments on IEMOCAP and CMU-MOSEI show that DGL-PCA consistently improves weighted F1 by 1.08–19.93% across all reproduced baselines. It achieves 70.02% and 83.91% weighted F1 on the IEMOCAP 6-way and 4-way settings, respectively, and 44.93% and 84.01% weighted F1 on the CMU-MOSEI 7-way and 2-way settings, respectively. Utterance-masking results demonstrated the robustness of the proposed method under dynamic emotional changes. In a 79-utterance IEMOCAP dialogue with 39 adjacent emotion transitions and up to nine transitions within a 15-utterance window, average weighted F1 slightly decreases by 1.87% in the case of any missing utterance, indicating long-context prediction stability under frequent emotion shifts. The proposed model paves the way for developing human-level emotion understanding capability of robots.

More from our Archive