Dynamic hypernode injection and graph attention fusion for multimodal sentiment analysis of behavioral signals
Wenfei Cao, Jian Xu, Yinghao Li, Huizi Yan, Zhuowei Hu, Wenxu Chen
Language use, vocal behavior, and facial activity provide complementary indicators of affective state, but multimodal models can be affected by noisy temporal observations and insufficient global guidance before cross-modal interaction. We developed a framework combining bidirectional gated recurrent unit encoders, attention pooling, dynamic hypernode injection, and graph attention fusion. Textual, acoustic, and visual sequences were mapped into a shared latent space and compressed into modality-level representations. A sample-dependent hypernode and a learnable static prior were then injected through gated residual connections before graph propagation. The model was evaluated on CMU-MOSI and CMU-MOSEI using five random seeds and validation-MAE checkpoint selection. On CMU-MOSI, the model obtained an MAE of