DOI: 10.3390/electronics15163624 ISSN: 2079-9292

Reliability-Aware Multi-Modal Sentiment Analysis Under Missing and Corrupted Modalities

Yubin Wu, Xianxun Zhu, Huilin Liu

Multi-modal sentiment analysis integrates linguistic, acoustic, and visual evidence, yet the reliability of these streams varies across samples because of missing observations, masking, measurement noise, and feature corruption. This paper presents a trainable reliability-aware evidential fusion framework that estimates not only sentiment predictions but also modality-specific evidence, predictive uncertainty, observable input quality, cross-modal disagreement, and normalized sample-dependent fusion weights. Each available modality is independently encoded and processed by an evidential classification head and a quality estimation head. Availability masks enforce exact exclusion of missing streams, while estimated quality, Dirichlet uncertainty, and Jensen–Shannon disagreement jointly regulate the contribution of each observed stream. The model is optimized end-to-end using fused classification, evidential regularization, clean–corrupted consistency, reliability-calibrated cross-modal alignment, and quality regression objectives. Experiments are conducted on both CMU-MOSI and CMU-MOSEI using their official speaker-independent splits. Binary classification follows the standard non-zero protocol, in which samples with sentiment score zero are excluded from Acc-2 and binary F1 evaluation; all labeled samples are retained for seven-class accuracy, mean absolute error, and correlation. The evaluation covers complete-input, every single- and double-modality missing pattern, graded and unseen corruption, combined missing-plus-corrupted conditions, calibration, selective prediction, statistical testing, and computational efficiency. All comparative values in the main tables are identified as local controlled adaptations under the common pipeline, while selected published reference values are reported separately to prevent provenance mixing. Across both datasets, the empirical results show that the proposed method preserves competitive complete-input performance while providing larger and more consistent gains as modality availability or integrity deteriorates.

More from our Archive