DOI: 10.1108/ajim-01-2026-0042 ISSN: 2050-3806

A hierarchical multistage semantic fusion framework for multimodal sentiment analysis

Guangyu Mu, Jiaxiu Dai, Yuanyuan Yue, Jiaxue Li

Purpose

Multimodal sentiment analysis aims to improve cross-modal fusion to understand sentiment better. Most existing methods rely on single-stage fusion or local cross-modal interactions, making it difficult to fully capture relationships among modalities, thereby limiting their sentiment representation capabilities and predictive performance.

Design/methodology/approach

This study proposes a Hierarchical Multistage Semantic Fusion (HMSF) framework. First, modality-specific encoding and unified projection are employed to achieve cross-modal alignment among textual, acoustic, and visual modalities. Next, a hierarchical multistage fusion structure is introduced to integrate multimodal information progressively. Finally, a gated contextual updating mechanism dynamically aggregates cross-sample contextual information to optimize fused representations. The model is trained with a unified objective for sentiment classification and continuous regression tasks.

Findings

Experimental results demonstrate that HMSF achieves competitive overall performance compared with representative existing methods on both the CMU-MOSEI and CH-SIMS datasets. On CMU-MOSEI, HMSF improves Acc-7 by approximately 4.60 percentage points over the average performance of the baseline methods. On CH-SIMS, HMSF improves the F1 score by approximately 1.44 percentage points over the average performance of the baseline methods.

Originality/value

By combining hierarchical fusion with gated contextual updating, HMSF enhances multimodal feature interaction and shows potential for applications such as online content analysis, public opinion monitoring, and user sentiment understanding.