Large Language Model-Assisted Distillation–Fusion Framework for Visual Emotion Recognition
Yujun Ma, Yunjie Zeng, Zhiyuan Chen, Zhiwei Ye, Wen Zhou, Chunli XiangVisual emotion recognition plays a critical role in human–computer interaction and mental health applications. Although existing Vision–Language Models (VLMs) alleviate the limitations of conventional vision models in high-level semantic understanding, they still face three main challenges: limited emotional semantic understanding, insufficient visual emotional perception capability, and high computational costs when deploying both models simultaneously. To address these issues, a large language model-assisted distillation–fusion framework (VERLADF) is proposed, which introduces emotion instruction data generated by GPT to fine-tune a VLM, thereby enhancing its emotional semantic understanding capability. Furthermore, we transfer the visual emotion discrimination knowledge of a conventional vision model into the VLM using a distillation module while keeping the VLM frozen during training, which reduces the computational costs. Following that, we design a fusion and prediction module that adaptively fuses predictions from the instruction-tuned VLM and the distillation module for final emotion recognition. The experimental results on the Abstract, ArtPhoto, Emotion6, and FI datasets demonstrate that VERLADF achieves recognition accuracies of 36.71%, 52.38%, 74.73%, and 79.69%, respectively, significantly outperforming many methods in the literature and demonstrating the effectiveness of the proposed framework.