Label-Aware Entropic Distributional Contrastive Alignment for Multimodal Sentiment Analysis
Mengyao Wang, Xiuyang Meng, Chunling WangMultimodal sentiment analysis integrates text, audio, and visual signals to infer affective states. However, sentiment information shared across modalities is often entangled with modality-specific variation, and existing representation learning methods do not fully exploit the relations encoded by continuous sentiment labels. Effective cross-view alignment therefore requires structured representations and a target that captures affective relations among samples. We propose a label-aware distributional contrastive alignment framework to address these requirements. Shared-private representation learning separates common sentiment information from modality-specific cues. The shared representations form audio–text and visual–text interaction views guided by a common textual reference, while the private representations retain complementary cues for prediction. Label-Aware GCA-UOT aligns the two sets of interaction views through batch-level distributional matching. Entropic unbalanced optimal transport provides a smooth, flexible coupling, and a label-aware target guides the allocation of matching mass according to paired-view correspondence, sentiment-intensity proximity, and polarity consistency. Experiments on CMU-MOSI, CMU-MOSEI, and CH-SIMS show consistent gains in fine-grained sentiment classification. Our method improves Acc-5 by 0.5–2.9 percentage points over the strongest baseline on each dataset and achieves the best or joint-best Acc-7 among the compared methods on CMU-MOSI and CMU-MOSEI.