Multimodal Sentiment and Emotion Detection in Social Media: Methods, Modalities, and Application Perspectives
Munmun Kakkar, Hemant PatidarThe growing popularity of social media has produced large volumes of multimodal user-generated content in the form of text, images, and speech, creating new opportunities to study how users express sentiment and emotion. This paper presents a narrative review of multimodal sentiment and emotion detection methods designed for social media data. The review covers representation methods for the textual, visual, and speech modalities, together with the machine learning and deep learning models used to analyze each modality and to fuse multimodal representations. The literature is analyzed in terms of feature extraction methods, fusion methods, and application contexts such as mental health monitoring, disaster management, education, and healthcare. The review shows that several weaknesses persist: text-based models still dominate, heterogeneous modalities are poorly integrated, standardized multimodal datasets are lacking, and deep multimodal systems are difficult to interpret. To address these concerns, the paper consolidates methodological knowledge across modalities and frames a unified multimodal analysis perspective that focuses on combining complementary features and on application-specific issues. The review offers a reference framework to guide future research and development in social media sentiment and emotion analysis by consolidating approaches, identifying research gaps, and clarifying the practical requirements of emotion-aware multimodal systems.