Sentiment Analysis of Microblog Opinion with Multimodal Feature Fusion
Dong Chen, YuXin Qian, YanGuo Pan, ZhenLong Du, XiaoLi LiAs an important research direction in the field of natural language processing, multimodal sentiment analysis is an essential part of semantic analysis, natural language understanding and other downstream tasks. Aiming at the problem that the feature differences between different modes in the existing methods make it difficult to fuse and utilize the cross-feature information, this paper proposes a cross-modal feature fusion model based on cross-attention mechanism and self-attention mechanism (multi-modal feature fusion, MMFF). The method uses separated feature extraction networks to encode different modalities and obtain unimodal features. In terms of cross-modal feature fusion, the cross-attention mechanism and the self-attention mechanism are used to obtain inter-modal interaction features and contextual features, and performs cross-modal fusion of different unimodal features to obtain the weights and correlations of the different semantic parts of the multiple feature sequences, so as to enhance the information interaction between different modalities. Multiple experiments on customized microblog Traffic opinion sentiment dataset and CH-SIMS dataset show that the proposed method can effectively utilize multimodal feature, and the experimental results are superior to most baseline multimodal sentiment analysis algorithms, which can effectively improve the performance of multimodal sentiment analysis task.