DOI: 10.3390/electronics15153382 ISSN: 2079-9292

EELLM: An Emotion-Enhanced Large Language Model for Multimodal Emotion Perception in IoT-Enabled Smart Sensor Networks

Lijiao Yang, Ming Cao, Ting Yang, Jinting Liu, Bachong Ma

AI-empowered smart sensor networks are driving Internet of Things (IoT) systems from passive data acquisition toward human-centric dynamic perception. In such scenarios, multimodal emotion recognition is expected to infer subtle affective states from heterogeneous audio, visual, and textual sensing streams. However, existing multimodal emotion recognition and emotion-oriented multimodal large-language-model (MLLM) methods still face several limitations for fine-grained emotion perception. Temporal asynchrony weakens cross-modal correspondence, redundant or heterogeneous features blur emotion-discriminative cues, unreliable sensing streams reduce robustness, and directly injecting all multimodal tokens into a large language model increases decoding redundancy. To address these issues, this paper proposes an Emotion-Enhanced Large Language Model (EELLM), a unified multimodal collaborative interaction framework for emotion perception in IoT-enabled smart sensor networks. Specifically, EELLM employs a dynamic time warping (DTW)-based Cross-Modal Alignment Module (DCAM) to mitigate temporal inconsistency, a Gaussian maximum mean discrepancy (MMD)-based Multimodal Feature Interaction Module (GMFIM) to disentangle shared and private representations and suppress fusion redundancy, a Modality Reliability-Aware Gating module (MRG) to adaptively weight heterogeneous modalities, and an Emotion-Salient Token Compression strategy (ESTC) to retain emotion-discriminative prefix tokens before instruction-guided LLaMA decoding. Extensive experiments on the trimodal MER2023 and MER2024 benchmarks and the visual-only DFEW benchmark demonstrate the effectiveness of EELLM. Under the adopted evaluation settings, EELLM achieves an F1 score of 0.9068 on MER2023, an average score of 67.10 on MER2024, and a UAR of 68.41 on DFEW, presenting competitive emotion perception performance across trimodal and visual-only benchmark settings. In addition, EELLM improves the recognition accuracy of the underrepresented disgust category on DFEW to 24.03%, showing better class-balanced emotion perception. These results indicate that EELLM provides an effective and efficient solution for fine-grained multimodal emotion perception in resource-sensitive intelligent sensing scenarios.

More from our Archive