Group Emotion Recognition Using Hybrid‐Level Fusion Strategy
Jingfeng Deng, Xinxin Wu, Ziwei Cui, Xingzhi WangABSTRACT
Group Emotion Recognition (GER) plays a vital role in public safety by enabling proactive risk warnings and resource allocation. However, current approaches often rely solely on visual emotional cues, and their performance is limited by linear fusion mechanisms. To overcome these limitations, this paper introduces a novel multimodal fusion framework that leverages Kolmogorov‐Arnold Networks (KANs) to effectively integrate multisource information. Our approach begins by extracting emotional features from three distinct modalities: group image scenes are processed by ResNet‐50, facial regions are analysed using VGGFace and textual descriptions, generated by the Gemini 2.0 model, are encoded with ELECTRA‐base. These features are then fused through an innovative hybrid‐level strategy based on KAN, which leverages its superior nonlinear modelling capability to generate robust emotional representations. Extensive experiments on three benchmark datasets, GAF 2.0, GAF 3.0 and GroupEmoW, demonstrate the consistent superiority of our method, which achieves state‐of‐the‐art accuracies of 92.87%, 92.01% and 91.88%, respectively. The advantage of the KAN‐based fusion is further validated through a comparison against seven alternative fusion mechanisms, paired statistical significance tests and an interpretability analysis of the learned modality contributions; moreover, we show that an open‐source MLLM (Qwen3‐VL) can replace the proprietary model with only a modest accuracy trade‐off, improving reproducibility and deployability.