Multimodal Decoupling Chain of Thought for Affective Understanding
Wenrui Zhang, Hao Yang, Shaobo Liu, Shuiting Du, Yixuan Chen, Xingpu Wu, Xiaowei ZhangThe Chain-of-Thought (CoT) technique has significantly improved the affective reasoning capabilities of large language models (LLMs). However, existing multimodal CoT typically aggregate information from different modalities into mixed inputs, which allows the model to perform entangled reasoning. This paradigm often exhibits cross-modal information confusion and misalignment when dealing with complex sentiment samples with semantic conflicts. To address these challenges, we propose a Multimodal Decoupling Chain-of-Thought for affective understanding called MDCoT. Specifically, MDCoT structures LLM reasoning into three explicit and progressive steps: (1) Shared Sentiment Aggregation, which extracts common emotional baselines aligned across all modalities; (2) Private Sentiment Disentanglement, which isolates modality-specific emotional nuances, such as textual metaphors, facial micro-expressions, or audio tone shifts; and (3) Global Integrated Inference, which balances shared consensus and modal-private clues to make the final prediction. Extensive experiments across three affective tasks (i.e., sentiment analysis, sarcasm detection, intent recognition) demonstrate that MDCoT consistently outperforms state-of-the-art end-to-end models and existing prompting strategies across multiple MLLMs. Furthermore, ablation and robustness studies confirm that explicit disentanglement effectively prevents cross-modal conflict and improves decision stability.